DataLearner logo

GPT-5.2 Pro Benchmark Details

GPT-5.2 Pro currently shows benchmark results led by GPQA Diamond (33 / 461, score 93.20), HLE (55 / 568, score 50), FrontierMath - Tier 4 (9 / 80, score 31.30). This page also compares it with 2 competitor models and 2 predecessor or same-series models, including performance and pricing views when available. 1 source link is attached for reference.

Benchmark Results

GPT-5.2 Pro

Benchmark Results

Thinking
Tool usage
Internet

Abstract Generalization

7 evaluations
Benchmark / mode
Score
Rank/total
ARC-AGI-1
MediumTools
81.17
88 / 176
ARC-AGI-1
Thinking Mode
90.50
60 / 176
ARC-AGI-1
HighTools
85.67
84 / 176
ARC-AGI-1
Extra-High
90.50
60 / 176
ARC-AGI-2
MediumTools
38.47
89 / 164
ARC-AGI-2
Thinking Mode
54.16
76 / 164
54.16
76 / 164

Knowledge Exams

2 evaluations
Benchmark / mode
Score
Rank/total
HLE
Thinking Mode
36.60
165 / 568
HLE
Thinking ModeTools
50
55 / 568

Scientific Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Mode
93.20
33 / 461

Commonsense

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Extra-High
57.40
39 / 93

Mathematics

4 evaluations
Benchmark / mode
Score
Rank/total
FrontierMath v2
Extra-High
74
12 / 58
46
14 / 42
FrontierMath - Tier 4
Standard ModeToolsInternet
31.30
9 / 80
31.30
9 / 80

Fact Finding

2 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
Thinking ModeTools
77.90
23 / 58
BrowseComp
Extra-HighTools
77.90
23 / 58

Capability Indices

1 evaluations
Benchmark / mode
Score
Rank/total
ECI
unknown
155.41
23 / 167

Competitor Comparison

Benchmark scores for GPT-5.2 Pro compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

9 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGPT-5.2 ProCurrentOpus 4.5
ARC-AGI-1
Score (%)
Abstract Generalization
90.50Thinking Level · Extra High
80.00Extended Thinking
ARC-AGI-2
Score (% solved); cost per task (USD)
Abstract Generalization
54.16Thinking Level · High
37.64Extended Thinking
HLE
Accuracy
Knowledge Exams
50.00Thinking Enabled | Tools
43.20Extended Thinking | Tools
GPQA Diamond
Accuracy
Scientific Reasoning
93.20Thinking Enabled
87.00Extended Thinking
SimpleBench
Score (AVG@5)
Commonsense
57.40Thinking Level · Extra High
62.00Extended Thinking
FrontierMath - Tier 4
Accuracy
Mathematics
31.30Standard Mode | Tools
4.20Standard Mode
FrontierMath Tier 4 v2
Accuracy (verification_code)
Mathematics
46.00Thinking Level · Extra High
4.8832K
FrontierMath v2
Accuracy (verification_code)
Mathematics
74.00Thinking Level · Extra High
34.3932K
ECI
ECI score (capability index, higher is better)
Capability Indices
155.41Thinking Level · High
150.10Thinking Level · High

Standard API Pricing: GPT-5.2 Pro vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
GPT-5.2 Pro
OpenAI$21 / 1M tokens$168 / 1M tokens—
Opus 4.5
Anthropic$5 / 1M tokens$25 / 1M tokens—

Version History

How each version of the GPT-5.2 Pro series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

9 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGPT-5.2 ProCurrentGPT-5-Pro
ARC-AGI-1
Score (%)
Abstract Generalization
90.50Thinking Level · Extra High
70.17Thinking Enabled
ARC-AGI-2
Score (% solved); cost per task (USD)
Abstract Generalization
54.16Thinking Level · High
18.00Thinking Enabled
HLE
Accuracy
Knowledge Exams
50.00Thinking Enabled | Tools
42.00Thinking Enabled | Tools
GPQA Diamond
Accuracy
Scientific Reasoning
93.20Thinking Enabled
89.40Thinking Enabled | Tools
SimpleBench
Score (AVG@5)
Commonsense
57.40Thinking Level · Extra High
61.60Thinking Enabled
FrontierMath - Tier 4
Accuracy
Mathematics
31.30Standard Mode | Tools
14.60Standard Mode
FrontierMath Tier 4 v2
Accuracy (verification_code)
Mathematics
46.00Thinking Level · Extra High
19.51Thinking Level · High
FrontierMath v2
Accuracy (verification_code)
Mathematics
74.00Thinking Level · Extra High
55.79Thinking Level · High
ECI
ECI score (capability index, higher is better)
Capability Indices
155.41Thinking Level · High
150.26Thinking Level · High

Single-Benchmark Version Trend

Viewing: ARC-AGI-1 · Abstract Generalization

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the GPT-5.2 Pro Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
GPT-5.2 Pro
OpenAI$21 / 1M tokens$168 / 1M tokens—
GPT-5.1 Pro
OpenAI$15 / 1M tokens$120 / 1M tokens—
GPT-5-Pro
OpenAI$15 / 1M tokens$120 / 1M tokens—

Sources