DataLearner logo

GPT-5.2 Benchmark Details

GPT-5.2 currently shows benchmark results led by AIME2025 (1 / 215, score 100), MMMU (1 / 74, score 85.90), LiveCodeBench (14 / 251, score 89.40). This page also compares it with 2 competitor models and 2 predecessor or same-series models, including performance and pricing views when available. 2 source links are attached for reference.

Benchmark Results

GPT-5.2

Benchmark Results

Thinking
Tool usage
Internet
Parallel

Knowledge Exams

7 evaluations
Benchmark / mode
Score
Rank/total
MMLU
Extra-High
89.60
12 / 124
HLE
Standard Mode
8
425 / 568
HLE
unknown
27.80
233 / 568
HLE
Medium
26.70
240 / 568
HLE
Extra-High
37.70
151 / 568
HLE
Extra-High
34.50
181 / 568
HLE
Extra-HighToolsInternet
45.50
85 / 568

Abstract Generalization

10 evaluations
Benchmark / mode
Score
Rank/total
55.67
127 / 176
ARC-AGI-1
Medium
72.67
101 / 176
78.67
91 / 176
ARC-AGI-1
Extra-High
86.17
79 / 176
ARC-AGI-1
Deep Thinking Mode
90.50
60 / 176
9.72
114 / 164
ARC-AGI-2
Medium
26.67
101 / 164
43.33
85 / 164
ARC-AGI-2
Extra-High
52.91
79 / 164
ARC-AGI-2
Deep Thinking Mode
54.20
75 / 164

Scientific Reasoning

9 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
73.23
299 / 461
82.70
199 / 461
87.88
121 / 461
88.19
117 / 461
GPQA Diamond
Extra-High
90.30
82 / 461
GPQA Diamond
Deep Thinking Mode
93.20
33 / 461
CritPt
Standard Mode
0.60
172 / 204
CritPt
Medium
7.90
88 / 204
CritPt
Extra-High
11.60
69 / 204

Repository Engineering

3 evaluations
Benchmark / mode
Score
Rank/total
SWE-bench Verified
Extra-HighTools
80
17 / 116
74.60
2 / 8
SWE-Bench Pro - Public
Extra-HighTools
55.60
30 / 62

Algorithmic Coding

3 evaluations
Benchmark / mode
Score
Rank/total
LiveCodeBench
Standard Mode
66.90
104 / 251
89.40
14 / 251
LiveCodeBench
Extra-High
88.90
16 / 251

Mathematics

16 evaluations
Benchmark / mode
Score
Rank/total
AIME2025
Standard Mode
51
154 / 215
AIME2025
Medium
96.70
21 / 215
AIME2025
Extra-High
100
1 / 215
AIME2025
Extra-HighTools
100
1 / 215
AIME 2026
HighTools
98.33
5 / 30
HMMT Feb 2026
HighTools
96.97
3 / 10
FrontierMath v2
Extra-High
67.40
17 / 58
FrontierMath
Extra-HighTools
40.30
8 / 60
31.70
19 / 42
6.30
35 / 80
16.70
20 / 80
18.80
16 / 80
18.80
16 / 80
FrontierMath - Tier 4
Extra-HighTools
14.60
23 / 80
13.54
13 / 17

Writing

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1699.80
26 / 106

Visual Understanding

7 evaluations
Benchmark / mode
Score
Rank/total
MMMU
Extra-High
85.90
1 / 74
MMMU
Extra-HighTools
80.40
18 / 74
VPCT
High
67
4 / 24
VPCT
Extra-High
84
2 / 24
MMMU-Pro
Standard Mode
65.80
150 / 229
MMMU-Pro
Medium
74.60
93 / 229
MMMU-Pro
Thinking Enabled
80.40
39 / 229

Commonsense

1 evaluations
Benchmark / mode
Score
Rank/total
45.80
58 / 93

Service Workflows

8 evaluations
Benchmark / mode
Score
Rank/total
τ²-Bench - Telecom
Standard ModeTools
46.50
178 / 264
74.30
127 / 264
89.69
67 / 264
τ²-Bench - Telecom
Extra-HighTools
84.80
90 / 264
τ²-Bench
Extra-HighTools
82
12 / 44
SAGE
Extra-HighTools
49.27
15 / 64
τ³-Banking
Standard ModeTools
11.10
130 / 167
τ³-Banking
HighTools
32.22
53 / 167

Instruction Following

3 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Standard Mode
47.40
168 / 282
IF Bench
Medium
65.20
100 / 282
IF Bench
Extra-High
75.40
34 / 282

Fact Finding

2 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
Extra-HighToolsInternet
65.80
34 / 58
BrowseComp
Extra-HighTools
65.80
34 / 58

Cross-capability Suites

4 evaluations
Benchmark / mode
Score
Rank/total
LiveBench
Standard Mode
48.91
96 / 117
65.33
53 / 117
LiveBench
Medium
71.84
30 / 117
74.63
16 / 117

Agentic Development

3 evaluations
Benchmark / mode
Score
Rank/total
Terminal Bench Hard
Standard ModeTools
31.80
100 / 244
43.20
41 / 244
Terminal Bench Hard
Extra-HighTools
47
26 / 244

Cross-industry Work

2 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA
HighTools
70.90
3 / 15
GDPval-AA
Extra-HighTools
61
4 / 15

Long Reasoning

3 evaluations
Benchmark / mode
Score
Rank/total
AA-LCR
Standard Mode
45.70
148 / 174
AA-LCR
Medium
70.30
104 / 174
AA-LCR
Extra-High
82.70
23 / 174

Tool Orchestration

1 evaluations
Benchmark / mode
Score
Rank/total
MCP-Atlas
Extra-HighTools
67.60
34 / 44

Maintenance & Optimization

1 evaluations
Benchmark / mode
Score
Rank/total
GSO
HighTools
27.40
7 / 21

ML Engineering

4 evaluations
Benchmark / mode
Score
Rank/total
WeirdML v2
Standard ModeTools
49.60
37 / 52
WeirdML v2
LowTools
49.60
37 / 52
WeirdML v2
MediumTools
63.40
25 / 52
WeirdML v2
Extra-HighTools
72.20
16 / 52

Preference Arenas

2 evaluations
Benchmark / mode
Score
Rank/total
1394
30 / 35
1480
19 / 35

Capability Frontier Metrics

1 evaluations
Benchmark / mode
Score
Rank/total
352.25
3 / 22

Code Generation & Editing

1 evaluations
Benchmark / mode
Score
Rank/total
Vibe Code Bench v1.1
Extra-HighTools
53.50
33 / 60

Clinical Workflows

2 evaluations
Benchmark / mode
Score
Rank/total
MedScribe
Extra-HighTools
84.39
24 / 66
MedCode
Extra-HighTools
49.75
12 / 64

Capability Indices

1 evaluations
Benchmark / mode
Score
Rank/total
ECI
unknown
153.45
35 / 167

Memory & Persistence

1 evaluations
Benchmark / mode
Score
Rank/total
EBR-bench
Extra-HighTools
23.02
12 / 23

Competitor Comparison

Benchmark scores for GPT-5.2 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGPT-5.2CurrentGemini 3.0 Pro (Preview 11-2025)Opus 4.5
HLE
Accuracy
Knowledge Exams
45.50Deep Thinking Mode | Tools
45.80Thinking Level · High | Tools
43.20Extended Thinking | Tools
ARC-AGI-1
Score (%)
Abstract Generalization
90.50Deep Thinking Mode
87.50Thinking Enabled
80.00Extended Thinking
ARC-AGI-2
Score (% solved); cost per task (USD)
Abstract Generalization
54.20Deep Thinking Mode
45.10Thinking Enabled
37.64Extended Thinking
CritPt
Score
Scientific Reasoning
11.60Thinking Level · Extra High
9.10Thinking Level · High
4.60Thinking Enabled
GPQA Diamond
Accuracy
Scientific Reasoning
93.20Deep Thinking Mode
93.80Thinking Enabled
87.00Extended Thinking
SWE-bench Verified
Accuracy
Repository Engineering
80.00Thinking Level · Extra High | Tools
76.20Thinking Enabled
80.90Extended Thinking | Tools
LiveCodeBench
Pass @K
Algorithmic Coding
89.40Thinking Level · Medium
92.00Thinking Enabled
87.10Thinking Enabled
AIME 2026
Accuracy
Mathematics
98.33Thinking Level · High | Tools
90.60Thinking Enabled
93.30Extended Thinking
AIME2025
Accuracy
Mathematics
100.00Thinking Level · Extra High | Tools
95.70Thinking Level · High
91.30Thinking Enabled
FrontierMath
Accuracy
Mathematics
40.30Thinking Level · Extra High | Tools
38.00Thinking Enabled
20.70Extended Thinking
FrontierMath - Tier 4
Accuracy
Mathematics
18.80Thinking Level · Extra High
18.80Standard Mode
4.20Standard Mode
FrontierMath Tier 4 v2
Accuracy (verification_code)
Mathematics
31.70Thinking Level · Extra High
--
4.8832K
27 additional benchmarks remain in the chart above.

Standard API Pricing: GPT-5.2 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Gemini 3.0 Pro (Preview 11-2025): Base price applies to <= 200000
ModelSupplierStandard inputStandard outputBase price applies to
GPT-5.2
OpenAI$1.75 / 1M tokens$14 / 1M tokens—
Gemini 3.0 Pro (Preview 11-2025)
Google DeepMind$2 / 1M tokens$12 / 1M tokens<= 200000
Opus 4.5
Anthropic$5 / 1M tokens$25 / 1M tokens—

Version History

How each version of the GPT-5.2 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGPT-5.2CurrentGPT-5.1GPT-5
HLE
Accuracy
Knowledge Exams
45.50Deep Thinking Mode | Tools
42.70Thinking Level · High | Tools
35.20Thinking Enabled | Tools
ARC-AGI-1
Score (%)
Abstract Generalization
90.50Deep Thinking Mode
72.83Thinking Level · High
65.67Thinking Level · High
ARC-AGI-2
Score (% solved); cost per task (USD)
Abstract Generalization
54.20Deep Thinking Mode
17.64Thinking Level · High
9.86Thinking Level · High
CritPt
Score
Scientific Reasoning
11.60Thinking Level · Extra High
4.90Thinking Level · High
5.70Thinking Level · High
GPQA Diamond
Accuracy
Scientific Reasoning
93.20Deep Thinking Mode
88.10Thinking Enabled
87.30Thinking Enabled | Tools
IC SWE-Lancer (Diamond)
Pass @K
Repository Engineering
74.60Thinking Level · Extra High | Tools
69.70Thinking Level · High
--
SWE-Bench Pro - Public
Accuracy
Repository Engineering
55.60Thinking Level · Extra High | Tools
50.80Thinking Level · High
36.30Thinking Level · High
SWE-bench Verified
Accuracy
Repository Engineering
80.00Thinking Level · Extra High | Tools
76.30Thinking Level · High | Tools
72.80Thinking Level · High
LiveCodeBench
Pass @K
Algorithmic Coding
89.40Thinking Level · Medium
86.80Thinking Level · High
84.60Thinking Level · High
AIME2025
Accuracy
Mathematics
100.00Thinking Level · Extra High | Tools
94.17Thinking Level · High | Tools
99.60Thinking Enabled | Tools
FrontierMath
Accuracy
Mathematics
40.30Thinking Level · Extra High | Tools
26.70Thinking Level · High | Tools
26.30Thinking Level · High | Tools
FrontierMath - Tier 4
Accuracy
Mathematics
18.80Thinking Level · Extra High
12.50Thinking Level · High | Tools
12.50Thinking Level · High
28 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: HLE · Knowledge Exams

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the GPT-5.2 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
GPT-5.2
OpenAI$1.75 / 1M tokens$14 / 1M tokens—
GPT-5.1
OpenAI$1.25 / 1M tokens$10 / 1M tokens—
GPT-5
OpenAI$1.25 / 1M tokens$10 / 1M tokens—

Sources