DataLearner logo

GPT-5.4 Pro Benchmark Details

GPT-5.4 Pro currently shows benchmark results led by GPQA Diamond (1 / 226, score 94.60), HLE (6 / 181, score 58.70), FrontierMath (3 / 60, score 50). This page also compares it with 2 competitor models and 3 predecessor or same-series models, including performance and pricing views when available. 2 source links are attached for reference.

Benchmark Results

GPT-5.4 Pro

Benchmark Results

Thinking
Tool usage
Internet

General Knowledge

4 evaluations
Benchmark / mode
Score
Rank/total
94.50
7 / 68
83.30
8 / 62
HLE
High
42.70
57 / 181
HLE
HighTools
58.70
6 / 181

Other

2 evaluations
Benchmark / mode
Score
Rank/total
94.40
3 / 226
GPQA Diamond
Extra-High
94.60
1 / 226

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
74.10
6 / 67

Math and Reasoning

7 evaluations
Benchmark / mode
Score
Rank/total
FrontierMath v2
Extra-High
82.46
7 / 34
58.54
8 / 34
50
3 / 60
FrontierMath
Extra-High
50
3 / 60
FrontierMath - Tier 4
Standard ModeToolsInternet
37.50
5 / 80
37.50
5 / 80

AI Agent - Information Search

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
HighTools
89.30
4 / 54

Productivity Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA
HighTools
82
8 / 21

Competitor Comparison

Benchmark scores for GPT-5.4 Pro compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

11 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGPT-5.4 ProCurrentClaude Opus 4.6Gemini 3.1 Pro Preview
ARC-AGI
综合评估
94.50Thinking Level · High
92.00Extended Thinking
--
ARC-AGI-2
综合评估
83.30Thinking Level · High
66.30Extended Thinking
77.10Thinking Level · High
HLE
综合评估
58.70Thinking Level · High | Tools
53.00Extended Thinking | Tools
51.40Thinking Level · High | Tools
GPQA Diamond
科学与综合推理
94.60Thinking Level · Extra High
91.31Extended Thinking
94.30Thinking Level · High
SimpleBench
常识推理
74.10Thinking Level · High
67.60Standard Mode
79.60Standard Mode
FrontierMath
数学推理
50.00Thinking Level · Extra High
40.70Thinking Level · High
36.90Thinking Level · High
38.00Thinking Level · High
22.90Thinking Level · High
16.70Standard Mode
58.54Thinking Level · Extra High
26.83Thinking Level · High
--
FrontierMath v2
数学推理
82.46Thinking Level · Extra High
65.96Thinking Level · High
--
BrowseComp
AI Agent - 信息收集
89.30Thinking Level · High | Tools
84.00Thinking Enabled | Tools
85.90Thinking Level · High | Tools
GDPval-AA
生产力知识
82.00Thinking Level · High | Tools
1606.00Extended Thinking | Tools
--

Standard API Pricing: GPT-5.4 Pro vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

GPT-5.4 Pro: Base price applies to <= 272K
Claude Opus 4.6: Base price applies to <= 200K
Gemini 3.1 Pro Preview: Base price applies to <= 200K
ModelSupplierStandard inputStandard outputBase price applies to
GPT-5.4 Pro
OpenAI$30 / 1M tokens$180 / 1M tokens<= 272K
Claude Opus 4.6
Anthropic$5 / 1M tokens$25 / 1M tokens<= 200K
Gemini 3.1 Pro Preview
Google Deep Mind$2 / 1M tokens$12 / 1M tokens<= 200K

Version History

How each version of the GPT-5.4 Pro series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

9 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGPT-5.4 ProCurrentGPT-5.2 ProGPT-5-Pro
ARC-AGI
综合评估
94.50Thinking Level · High
90.50Thinking Enabled
70.20Thinking Enabled
ARC-AGI-2
综合评估
83.30Thinking Level · High
54.20Thinking Enabled
18.00Thinking Enabled
HLE
综合评估
58.70Thinking Level · High | Tools
50.00Thinking Enabled | Tools
42.00Thinking Enabled | Tools
GPQA Diamond
科学与综合推理
94.60Thinking Level · Extra High
93.20Thinking Enabled
89.40Thinking Enabled | Tools
SimpleBench
常识推理
74.10Thinking Level · High
57.40Thinking Level · Extra High
61.60Thinking Enabled
38.00Thinking Level · High
31.30Thinking Enabled
14.60Thinking Enabled
58.54Thinking Level · Extra High
46.00Thinking Level · Extra High
19.51Thinking Level · High
FrontierMath v2
数学推理
82.46Thinking Level · Extra High
74.00Thinking Level · Extra High
55.79Thinking Level · High
BrowseComp
AI Agent - 信息收集
89.30Thinking Level · High | Tools
77.90Thinking Enabled | Tools
--

Single-Benchmark Version Trend

Viewing: ARC-AGI · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the GPT-5.4 Pro Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

GPT-5.4 Pro: Base price applies to <= 272K
ModelSupplierStandard inputStandard outputBase price applies to
GPT-5.4 Pro
OpenAI$30 / 1M tokens$180 / 1M tokens<= 272K
GPT-5.2 Pro
OpenAI$21 / 1M tokens$168 / 1M tokens
GPT-5.1 Pro
OpenAI$15 / 1M tokens$120 / 1M tokens
GPT-5-Pro
OpenAI$15 / 1M tokens$120 / 1M tokens

Sources