DataLearner logo

GPT-5.5 Pro Benchmark Details

GPT-5.5 Pro currently shows benchmark results led by FrontierMath - Tier 4 (1 / 80, score 39.60), FrontierMath (1 / 60, score 52.40), FrontierMath v2 (2 / 58, score 87.72). This page also compares it with 3 competitor models and 3 predecessor or same-series models, including performance and pricing views when available. 1 source link is attached for reference.

Benchmark Results

GPT-5.5 Pro

Benchmark Results

Thinking
Tool usage
Internet

General Knowledge

6 evaluations
Benchmark / mode
Score
Rank/total
96.50
9 / 92
ARC-AGI-1
Extra-High
95
17 / 92
84.60
15 / 85
ARC-AGI-2
Extra-High
84.20
18 / 85
HLE
Extra-High
43.10
63 / 197
HLE
Extra-HighTools
57.20
14 / 197

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Extra-High
93.92
13 / 274

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Standard Mode
76.90
9 / 94

Math and Reasoning

7 evaluations
Benchmark / mode
Score
Rank/total
88.10
1 / 19
FrontierMath v2
Extra-High
87.72
2 / 58
78.05
5 / 42
FrontierMath
Extra-HighTools
52.40
1 / 60
39.60
1 / 80
39.60
1 / 80
FrontierMath - Tier 4
Extra-HighTools
39.60
1 / 80

AI Agent - Information Search

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
Extra-HighToolsInternet
90.10
5 / 57

Productivity Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA
Extra-High
82.30
7 / 21

Competitor Comparison

Benchmark scores for GPT-5.5 Pro compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

10 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGPT-5.5 ProCurrentClaude Mythos PreviewGemini 3.1 Pro PreviewOpus 4.7
ARC-AGI-1
Score (%)
综合评估
96.50Thinking Level · High
--
--
93.50Thinking Level · High
ARC-AGI-2
Accuracy and cost per task
综合评估
84.60Thinking Level · High
--
77.10Thinking Level · High
75.80Thinking Level · High
HLE
Accuracy
综合评估
57.20Thinking Level · Extra High | Tools
64.70Extended Thinking | Tools
51.40Thinking Level · High | Tools
54.70Extended Thinking | Tools
GPQA Diamond
Accuracy
科学与综合推理
93.92Thinking Level · Extra High
94.60Extended Thinking
94.30Thinking Level · High
94.20Extended Thinking
SimpleBench
Score (AVG@5)
常识推理
76.90Standard Mode
--
79.60Standard Mode
61.70Standard Mode
FrontierMath
Accuracy
数学推理
52.40Thinking Level · Extra High | Tools
--
36.90Thinking Level · High
43.80Thinking Level · Extra High
FrontierMath - Tier 4
Accuracy
数学推理
39.60Thinking Level · Extra High | Tools
--
16.70Standard Mode
22.90Thinking Level · Extra High
FrontierMath Tier 4 v2
Accuracy (verification_code)
数学推理
78.05Thinking Level · Extra High
--
--
31.71Thinking Level · High
FrontierMath v2
Accuracy (verification_code)
数学推理
87.72Thinking Level · Extra High
--
--
70.18Thinking Level · High
BrowseComp
Accuracy
AI Agent - 信息收集
90.10Deep Thinking Mode | Tools
84.90Extended Thinking | Tools
85.90Thinking Level · High | Tools
79.30Extended Thinking | Tools

Standard API Pricing: GPT-5.5 Pro vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Gemini 3.1 Pro Preview: Base price applies to <= 200K
ModelSupplierStandard inputStandard outputBase price applies to
GPT-5.5 Pro
OpenAI$30 / 1M tokens$180 / 1M tokens
Claude Mythos Preview
Anthropic$25 / 1M tokens$125 / 1M tokens
Gemini 3.1 Pro Preview
Google Deep Mind$2 / 1M tokens$12 / 1M tokens<= 200K
Opus 4.7
Anthropic$5 / 1M tokens$25 / 1M tokens

Version History

How each version of the GPT-5.5 Pro series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

11 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGPT-5.5 ProCurrentGPT-5.4 ProGPT-5.2 Pro
ARC-AGI-1
Score (%)
综合评估
96.50Thinking Level · High
94.50Thinking Level · High
90.50Thinking Enabled
ARC-AGI-2
Accuracy and cost per task
综合评估
84.60Thinking Level · High
83.30Thinking Level · High
54.20Thinking Enabled
HLE
Accuracy
综合评估
57.20Thinking Level · Extra High | Tools
58.70Thinking Level · High | Tools
50.00Thinking Enabled | Tools
GPQA Diamond
Accuracy
科学与综合推理
93.92Thinking Level · Extra High
94.60Thinking Level · Extra High
93.20Thinking Enabled
SimpleBench
Score (AVG@5)
常识推理
76.90Standard Mode
74.10Thinking Level · High
57.40Thinking Level · Extra High
FrontierMath
Accuracy
数学推理
52.40Thinking Level · Extra High | Tools
50.00Thinking Level · Extra High
--
FrontierMath - Tier 4
Accuracy
数学推理
39.60Thinking Level · Extra High | Tools
38.00Thinking Level · High
31.30Standard Mode | Tools
FrontierMath Tier 4 v2
Accuracy (verification_code)
数学推理
78.05Thinking Level · Extra High
58.54Thinking Level · Extra High
46.00Thinking Level · Extra High
FrontierMath v2
Accuracy (verification_code)
数学推理
87.72Thinking Level · Extra High
82.46Thinking Level · Extra High
74.00Thinking Level · Extra High
BrowseComp
Accuracy
AI Agent - 信息收集
90.10Deep Thinking Mode | Tools
89.30Thinking Level · High | Tools
77.90Thinking Level · Extra High | Tools
GDPval-AA
Accuracy
生产力知识
82.30Thinking Level · Extra High
82.00Thinking Level · High | Tools
--

Single-Benchmark Version Trend

Viewing: ARC-AGI-1 · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the GPT-5.5 Pro Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
GPT-5.5 Pro
OpenAI$30 / 1M tokens$180 / 1M tokens
GPT-5.4 Pro
OpenAI$30 / 1M tokens$180 / 1M tokens
GPT-5.2 Pro
OpenAI$21 / 1M tokens$168 / 1M tokens
GPT-5.1 Pro
OpenAI$15 / 1M tokens$120 / 1M tokens

Sources