DataLearner logo

GPT-5.2 Benchmark Details

GPT-5.2 currently shows benchmark results led by AIME2025 (1 / 106, score 100), MMMU (1 / 28, score 85.90), GPQA Diamond (20 / 271, score 93.20). This page also compares it with 2 competitor models and 2 predecessor or same-series models, including performance and pricing views when available. 2 source links are attached for reference.

Benchmark Results

GPT-5.2

Benchmark Results

Thinking
Tool usage
Internet
Parallel

General Knowledge

17 evaluations
Benchmark / mode
Score
Rank/total
55.70
67 / 91
ARC-AGI-1
Medium
72.70
52 / 91
78.70
46 / 91
ARC-AGI-1
Extra-High
86.20
41 / 91
ARC-AGI-1
Deep Thinking Mode
90.50
32 / 91
MMLU
Extra-High
89.60
12 / 124
LiveBench
Standard Mode
48.91
94 / 115
65.33
53 / 115
LiveBench
Medium
71.84
31 / 115
74.84
19 / 115
9.70
64 / 85
ARC-AGI-2
Medium
26.70
57 / 85
43.30
48 / 85
ARC-AGI-2
Extra-High
52.90
45 / 85
ARC-AGI-2
Deep Thinking Mode
54.20
43 / 85
HLE
Extra-High
34.50
89 / 190
HLE
Extra-HighToolsInternet
45.50
49 / 190

Other

6 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
73.23
181 / 271
82.70
125 / 271
87.88
72 / 271
88.19
67 / 271
GPQA Diamond
Extra-High
92.40
26 / 271
GPQA Diamond
Deep Thinking Mode
93.20
20 / 271

Coding and Software Engineer

10 evaluations
Benchmark / mode
Score
Rank/total
1394
30 / 35
1480
19 / 35
SWE-bench Verified
Extra-HighTools
80
17 / 115
IC SWE-Lancer(Diamond)
Extra-HighTools
74.60
2 / 8
WeirdML v2
Standard ModeTools
49.60
37 / 52
WeirdML v2
LowTools
49.60
37 / 52
WeirdML v2
MediumTools
63.40
25 / 52
WeirdML v2
Extra-HighTools
72.20
16 / 52
SWE-Bench Pro - Public
Extra-HighTools
55.60
29 / 60
GSO
HighTools
27.40
7 / 21

Math and Reasoning

9 evaluations
Benchmark / mode
Score
Rank/total
AIME2025
Extra-High
100
1 / 106
FrontierMath v2
Extra-High
67.40
17 / 58
FrontierMath
Extra-HighTools
40.30
8 / 60
31.70
18 / 41
6.30
35 / 80
16.70
20 / 80
18.80
16 / 80
18.80
16 / 80
FrontierMath - Tier 4
Extra-HighTools
14.60
23 / 80

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1698.90
22 / 99

Multimodal Understanding

4 evaluations
Benchmark / mode
Score
Rank/total
MMMU
Extra-High
85.90
1 / 28
MMMU
Extra-HighTools
80.40
12 / 28
VPCT
High
67
4 / 24
VPCT
Extra-High
84
2 / 24

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
45.80
56 / 92

Agent Level Benchmark

3 evaluations
Benchmark / mode
Score
Rank/total
352.25
3 / 22
τ²-Bench - Telecom
Extra-HighTools
98.70
4 / 35
τ²-Bench
Extra-HighTools
82
12 / 44

AI Agent - Information Search

2 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
Extra-HighToolsInternet
65.80
32 / 56
BrowseComp
Extra-HighTools
65.80
32 / 56

Productivity Knowledge

2 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA
HighTools
70.90
9 / 21
GDPval-AA
Extra-HighTools
61
10 / 21

AI Agent - Tool Usage

1 evaluations
Benchmark / mode
Score
Rank/total
MCP-Atlas
Extra-HighTools
67.60
31 / 41

Competitor Comparison

Benchmark scores for GPT-5.2 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGPT-5.2CurrentGemini 3.0 Pro (Preview 11-2025)Opus 4.5
ARC-AGI-1
综合评估
90.50Deep Thinking Mode
87.50Thinking Enabled
80.00Extended Thinking
ARC-AGI-2
综合评估
54.20Deep Thinking Mode
45.10Thinking Enabled
37.60Extended Thinking
HLE
综合评估
45.50Deep Thinking Mode | Tools
45.80Thinking Level · High | Tools
43.20Extended Thinking | Tools
LiveBench
综合评估
74.84Thinking Level · High
73.39Thinking Level · High
75.9664K
GPQA Diamond
科学与综合推理
93.20Deep Thinking Mode
93.80Thinking Enabled
87.00Extended Thinking
GSO
编程与软件工程
27.40Thinking Level · High | Tools
--
26.50Standard Mode | Tools
SWE-bench Verified
编程与软件工程
80.00Thinking Level · Extra High | Tools
76.20Thinking Enabled
80.90Extended Thinking | Tools
Text Arena (Coding)
编程与软件工程
1480.00Thinking Level · High
--
1512.0032K
WeirdML v2
编程与软件工程
72.20Thinking Level · Extra High | Tools
--
63.7016K | Tools
AIME2025
数学推理
100.00Thinking Level · Extra High
95.00Thinking Enabled
--
FrontierMath
数学推理
40.30Thinking Level · Extra High | Tools
38.00Thinking Enabled
20.70Extended Thinking
18.80Thinking Level · Extra High
18.80Standard Mode
4.20Standard Mode
12 additional benchmarks remain in the chart above.

Standard API Pricing: GPT-5.2 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Gemini 3.0 Pro (Preview 11-2025): Base price applies to <= 200000
ModelSupplierStandard inputStandard outputBase price applies to
GPT-5.2
Facebook AI研究实验室$1.75 / 1M tokens$14 / 1M tokens
Gemini 3.0 Pro (Preview 11-2025)
Google Deep Mind$2 / 1M tokens$12 / 1M tokens<= 200000
Opus 4.5
Facebook AI研究实验室$5 / 1M tokens$25 / 1M tokens

Version History

How each version of the GPT-5.2 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGPT-5.2CurrentGPT-5.1GPT-5
ARC-AGI-1
综合评估
90.50Deep Thinking Mode
72.80Thinking Level · High
65.70Thinking Level · High
ARC-AGI-2
综合评估
54.20Deep Thinking Mode
17.60Thinking Level · High
9.90Thinking Level · High
HLE
综合评估
45.50Deep Thinking Mode | Tools
42.70Thinking Level · High | Tools
35.20Thinking Enabled | Tools
LiveBench
综合评估
74.84Thinking Level · High
72.04Thinking Level · High
--
GPQA Diamond
科学与综合推理
93.20Deep Thinking Mode
88.10Thinking Level · High
87.30Thinking Enabled | Tools
GSO
编程与软件工程
27.40Thinking Level · High | Tools
13.70Thinking Level · High | Tools
6.90Thinking Level · High | Tools
IC SWE-Lancer(Diamond)
编程与软件工程
74.60Thinking Level · Extra High | Tools
69.70Thinking Level · High
--
SWE-Bench Pro - Public
编程与软件工程
55.60Thinking Level · Extra High | Tools
50.80Thinking Level · High
36.30Thinking Level · High
SWE-bench Verified
编程与软件工程
80.00Thinking Level · Extra High | Tools
76.30Thinking Level · High | Tools
72.80Thinking Level · High
Text Arena (Coding)
编程与软件工程
1480.00Thinking Level · High
1387.00Thinking Level · Medium
1395.00Thinking Level · Medium
WeirdML v2
编程与软件工程
72.20Thinking Level · Extra High | Tools
60.77Thinking Level · High | Tools
60.70Thinking Level · High | Tools
AIME2025
数学推理
100.00Thinking Level · Extra High
94.00Thinking Level · High
99.60Thinking Enabled | Tools
13 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: ARC-AGI-1 · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the GPT-5.2 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
GPT-5.2
Facebook AI研究实验室$1.75 / 1M tokens$14 / 1M tokens
GPT-5.1
OpenAI$1.25 / 1M tokens$10 / 1M tokens
GPT-5
OpenAI$1.25 / 1M tokens$10 / 1M tokens

Sources