DataLearner logo

GPT-5.5 Benchmark Details

GPT-5.5 currently shows benchmark results led by LiveBench (1 / 117, score 79.91), Terminal Bench 2.0 (1 / 48, score 82.70), FrontierMath (2 / 60, score 51.70). This page also compares it with 3 competitor models and 3 predecessor or same-series models, including performance and pricing views when available. 2 source links are attached for reference.

Benchmark Results

GPT-5.5

Benchmark Results

Thinking
Tool usage
Internet

General Knowledge

15 evaluations
Benchmark / mode
Score
Rank/total
76.20
48 / 92
ARC-AGI-1
Medium
92.20
27 / 92
94.50
19 / 92
ARC-AGI-1
Extra-High
95
17 / 92
33.30
53 / 85
ARC-AGI-2
Medium
70.40
29 / 85
85
13 / 85
ARC-AGI-2
Extra-High
85
13 / 85
LiveBench
Standard Mode
72.15
28 / 117
LiveBench
Medium
68.66
44 / 117
78.75
4 / 117
LiveBench
Deep Thinking Mode
79.91
1 / 117
HLE
High
41.40
71 / 197
HLE
HighTools
52.20
28 / 197

Other

4 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
77.27
164 / 274
90.66
40 / 274
93.60
15 / 274
GPQA Diamond
Extra-High
94
11 / 274

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1843.50
12 / 106

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Standard Mode
69
18 / 94

Math and Reasoning

6 evaluations
Benchmark / mode
Score
Rank/total
FrontierMath v2
Extra-High
85.26
6 / 58
72.50
7 / 42
71.90
3 / 19
FrontierMath
HighTools
51.70
2 / 60
35.40
7 / 80
35.40
7 / 80

Coding and Software Engineer

8 evaluations
Benchmark / mode
Score
Rank/total
1478.93
21 / 35
1504.74
18 / 35
WeirdML v2
Standard ModeTools
67.15
22 / 52
WeirdML v2
HighTools
83.90
6 / 52
WeirdML v2
Extra-HighTools
84.91
5 / 52
DeepSWE
Extra-HighTools
67
14 / 38
58.60
18 / 62
GSO
Extra-HighTools
40.20
4 / 21

Agent Level Benchmark

1 evaluations
Benchmark / mode
Score
Rank/total
98
5 / 36

AI Agent - Information Search

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
HighToolsInternet
84.40
10 / 57

AI Agent - Tool Usage

5 evaluations
Benchmark / mode
Score
Rank/total
83.40
20 / 53
Terminal-Bench 2.1
Extra-HighTools
83.10
22 / 53
82.70
1 / 48
78.70
9 / 26
MCP-Atlas
Extra-HighTools
75.30
23 / 41

Text Embedding

5 evaluations
Benchmark / mode
Score
Rank/total
Context Arena
Standard Mode
41.93
98 / 126
87.46
20 / 126
90.79
13 / 126
92.01
12 / 126
Context Arena
Extra-High
94.18
6 / 126

Productivity Knowledge

2 evaluations
Benchmark / mode
Score
Rank/total
1769
2 / 21
AA-AnalystAgent
Extra-HighToolsInternet
50
3 / 12

Competitor Comparison

Benchmark scores for GPT-5.5 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGPT-5.5CurrentOpus 4.7Claude Mythos PreviewGemini 3.1 Pro Preview
ARC-AGI-1
Score (%)
综合评估
95.00Thinking Level · Extra High
93.50Thinking Level · High
--
--
ARC-AGI-2
Accuracy and cost per task
综合评估
85.00Thinking Level · Extra High
75.80Thinking Level · High
--
77.10Thinking Level · High
HLE
Accuracy
综合评估
52.20Thinking Level · High | Tools
54.70Extended Thinking | Tools
64.70Extended Thinking | Tools
51.40Thinking Level · High | Tools
LiveBench
Accuracy
综合评估
79.91Deep Thinking Mode
76.53Deep Thinking Mode
--
77.13Thinking Level · High
GPQA Diamond
Accuracy
科学与综合推理
94.00Thinking Level · Extra High
94.20Extended Thinking
94.60Extended Thinking
94.30Thinking Level · High
Creative Writing
Elo、大模型评判两两对战
写作和创作
1843.50Standard Mode
1907.10Standard Mode
--
1488.90Standard Mode
SimpleBench
Score (AVG@5)
常识推理
69.00Standard Mode
61.70Standard Mode
--
79.60Standard Mode
FrontierMath
Accuracy
数学推理
51.70Thinking Level · High | Tools
43.80Thinking Level · Extra High
--
36.90Thinking Level · High
FrontierMath - Tier 4
Accuracy
数学推理
35.40Thinking Level · Extra High
22.90Thinking Level · Extra High
--
16.70Standard Mode
FrontierMath Tier 4 v2
Accuracy (verification_code)
数学推理
72.50Thinking Level · Extra High
31.71Thinking Level · High
--
--
FrontierMath v2
Accuracy (verification_code)
数学推理
85.26Thinking Level · Extra High
70.18Thinking Level · High
--
--
DeepSWE
Pass@1 (DeepSWE v1.1)
编程与软件工程
67.00Thinking Level · Extra High | Tools
--
--
12.00Thinking Level · High | Tools
11 additional benchmarks remain in the chart above.

Standard API Pricing: GPT-5.5 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Gemini 3.1 Pro Preview: Base price applies to <= 200K
ModelSupplierStandard inputStandard outputBase price applies to
GPT-5.5
OpenAI$5 / 1M tokens$30 / 1M tokens
Opus 4.7
Anthropic$5 / 1M tokens$25 / 1M tokens
Claude Mythos Preview
Anthropic$25 / 1M tokens$125 / 1M tokens
Gemini 3.1 Pro Preview
Google Deep Mind$2 / 1M tokens$12 / 1M tokens<= 200K

Version History

How each version of the GPT-5.5 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGPT-5.5CurrentGPT-5.4GPT-5.2GPT-5.1
ARC-AGI-1
Score (%)
综合评估
95.00Thinking Level · Extra High
93.70Thinking Level · Extra High
90.50Deep Thinking Mode
72.80Thinking Level · High
ARC-AGI-2
Accuracy and cost per task
综合评估
85.00Thinking Level · Extra High
77.10Standard Mode
54.20Deep Thinking Mode
17.60Thinking Level · High
HLE
Accuracy
综合评估
52.20Thinking Level · High | Tools
52.10Thinking Level · Extra High | Tools
45.50Deep Thinking Mode | Tools
42.70Thinking Level · High | Tools
LiveBench
Accuracy
综合评估
79.91Deep Thinking Mode
77.97Deep Thinking Mode
74.63Thinking Level · High
72.04Thinking Level · High
GPQA Diamond
Accuracy
科学与综合推理
94.00Thinking Level · Extra High
92.80Thinking Level · Extra High
93.20Deep Thinking Mode
88.10Thinking Level · High
Creative Writing
Elo、大模型评判两两对战
写作和创作
1843.50Standard Mode
1835.60Standard Mode
1699.80Standard Mode
--
SimpleBench
Score (AVG@5)
常识推理
69.00Standard Mode
--
45.80Thinking Level · High
53.20Thinking Level · High
FrontierMath
Accuracy
数学推理
51.70Thinking Level · High | Tools
47.60Thinking Level · Extra High
40.30Thinking Level · Extra High | Tools
26.70Thinking Level · High | Tools
FrontierMath - Tier 4
Accuracy
数学推理
35.40Thinking Level · Extra High
27.10Thinking Level · Extra High
18.80Thinking Level · Extra High
12.50Thinking Level · High | Tools
FrontierMath Tier 4 v2
Accuracy (verification_code)
数学推理
72.50Thinking Level · Extra High
49.00Thinking Level · Extra High
31.70Thinking Level · Extra High
--
FrontierMath v2
Accuracy (verification_code)
数学推理
85.26Thinking Level · Extra High
78.60Thinking Level · Extra High
67.40Thinking Level · Extra High
--
DeepSWE
Pass@1 (DeepSWE v1.1)
编程与软件工程
67.00Thinking Level · Extra High | Tools
52.00Thinking Level · Extra High | Tools
--
--
11 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: ARC-AGI-1 · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the GPT-5.5 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
GPT-5.5
OpenAI$5 / 1M tokens$30 / 1M tokens
GPT-5.4
OpenAI$2.5 / 1M tokens$15 / 1M tokens
GPT-5.2
Facebook AI研究实验室$1.75 / 1M tokens$14 / 1M tokens
GPT-5.1
OpenAI$1.25 / 1M tokens$10 / 1M tokens

Sources