DataLearner logo

GPT-5.4 Benchmark Details

GPT-5.4 currently shows benchmark results led by Pinch Bench (1 / 38, score 90.50), LiveBench (5 / 117, score 77.97), Terminal Bench Hard (11 / 244, score 57.60). This page also compares it with 2 competitor models and 2 predecessor or same-series models, including performance and pricing views when available. 2 source links are attached for reference.

Benchmark Results

GPT-5.4

Benchmark Results

Thinking
Tool usage

Abstract Generalization

11 evaluations
Benchmark / mode
Score
Rank/total
ARC-AGI-1
Standard Mode
93.67
40 / 176
68.17
104 / 176
ARC-AGI-1
Medium
86.17
79 / 176
ARC-AGI-1
HighTools
92.67
44 / 176
ARC-AGI-1
Extra-High
93.67
40 / 176
ARC-AGI-2
Standard Mode
77.10
38 / 164
29.17
99 / 164
ARC-AGI-2
Medium
55.42
73 / 164
ARC-AGI-2
HighTools
67.50
52 / 164
ARC-AGI-2
Extra-High
73.95
43 / 164

Knowledge Exams

5 evaluations
Benchmark / mode
Score
Rank/total
HLE
Standard Mode
11.30
380 / 568
HLE
Low
30.80
197 / 568
HLE
Extra-High
43.70
96 / 568
HLE
Extra-High
36.24
167 / 568
HLE
Extra-HighTools
52.10
45 / 568

Scientific Reasoning

8 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
74.80
285 / 461
84.85
160 / 461
88.89
104 / 461
89.90
89 / 461
GPQA Diamond
Extra-High
92
53 / 461
CritPt
Standard Mode
0.60
172 / 204
7.40
89 / 204
CritPt
Extra-High
23.40
30 / 204

Writing

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1835.60
14 / 106

Mathematics

7 evaluations
Benchmark / mode
Score
Rank/total
AIME 2026
Extra-HighTools
99.17
4 / 30
HMMT Feb 2026
Extra-HighTools
97.73
2 / 10
FrontierMath v2
Extra-High
78.60
10 / 58
MathArena Apex
Extra-HighTools
54.17
8 / 17
49
12 / 42
FrontierMath
Extra-High
47.60
5 / 60
27.10
11 / 80

Repository Engineering

3 evaluations
Benchmark / mode
Score
Rank/total
57.70
22 / 62
DeepSWE
Extra-HighTools
51.77
64 / 91
43.40
2 / 3

Service Workflows

5 evaluations
Benchmark / mode
Score
Rank/total
τ²-Bench - Telecom
Standard ModeTools
64.30
151 / 264
74.60
126 / 264
τ²-Bench - Telecom
Extra-HighTools
87.10
74 / 264
SAGE
Extra-HighTools
43.31
37 / 64
τ³-Banking
Extra-HighTools
39.43
36 / 167

Instruction Following

3 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Standard Mode
48.40
164 / 282
65.90
98 / 282
IF Bench
Extra-High
73.90
42 / 282

Fact Finding

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
Extra-HighTools
82.70
18 / 58

Cross-capability Suites

2 evaluations
Benchmark / mode
Score
Rank/total
75.07
12 / 117
LiveBench
Deep Thinking Mode
77.97
5 / 117

Agentic Development

5 evaluations
Benchmark / mode
Score
Rank/total
Terminal-Bench 2.1
Extra-HighTools
78.30
71 / 199
Terminal Bench 2.0
Extra-HighTools
75.10
4 / 48
Terminal Bench Hard
Standard ModeTools
37.90
59 / 244
43.20
41 / 244
Terminal Bench Hard
Extra-HighTools
57.60
11 / 244

Memory & Persistence

6 evaluations
Benchmark / mode
Score
Rank/total
Context Arena
Standard Mode
32.79
112 / 126
74.16
52 / 126
77.84
45 / 126
82.99
28 / 126
Context Arena
Extra-High
86.15
22 / 126
EBR-bench
Extra-HighTools
25.40
11 / 23

Long Reasoning

3 evaluations
Benchmark / mode
Score
Rank/total
AA-LCR
Standard Mode
58.30
134 / 174
76.70
81 / 174
AA-LCR
Extra-High
82
28 / 174

Desktop Workflows

1 evaluations
Benchmark / mode
Score
Rank/total
OSWorld-Verified
Extra-HighTools
75
15 / 28

Tool Orchestration

4 evaluations
Benchmark / mode
Score
Rank/total
Claw Bench
Thinking EnabledTools
92.70
3 / 29
Pinch Bench
Thinking EnabledTools
90.50
1 / 38
75.70
17 / 45
MCP-Atlas
Extra-HighTools
70.60
29 / 44

Cross-industry Work

1 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA v2
Extra-HighTools
1307
65 / 110

Visual Understanding

5 evaluations
Benchmark / mode
Score
Rank/total
MMMU-Pro
Standard Mode
70.60
125 / 229
78
64 / 229
MMMU-Pro
Thinking Enabled
81.20
32 / 229
MMMU-Pro
Thinking EnabledTools
82.10
25 / 229
MMMU-Pro
Extra-High
78.40
61 / 229

Maintenance & Optimization

3 evaluations
Benchmark / mode
Score
Rank/total
Code Migration
Extra-HighTools
34.98
25 / 43
GSO
HighTools
25.49
9 / 21
GSO
Extra-HighTools
31.37
6 / 21

ML Engineering

2 evaluations
Benchmark / mode
Score
Rank/total
WeirdML v2
Standard ModeTools
57.44
31 / 52
WeirdML v2
Extra-HighTools
77.70
12 / 52

Preference Arenas

2 evaluations
Benchmark / mode
Score
Rank/total
1437.34
26 / 35
1457.23
24 / 35

Capability Frontier Metrics

1 evaluations
Benchmark / mode
Score
Rank/total
341.73
5 / 22

Honesty & Factuality

1 evaluations
Benchmark / mode
Score
Rank/total
5.78
34 / 47

Code Generation & Editing

1 evaluations
Benchmark / mode
Score
Rank/total
Vibe Code Bench v1.1
Extra-HighTools
48.47
37 / 60

Clinical Workflows

2 evaluations
Benchmark / mode
Score
Rank/total
MedScribe
Extra-HighTools
77.55
46 / 66
MedCode
Extra-HighTools
41.29
41 / 64

Legal

1 evaluations
Benchmark / mode
Score
Rank/total
0
36 / 43

Capability Indices

1 evaluations
Benchmark / mode
Score
Rank/total
ECI
unknown
156.92
14 / 167

Competitor Comparison

Benchmark scores for GPT-5.4 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGPT-5.4CurrentGemini 3.1 Pro PreviewClaude Opus 4.6
ARC-AGI-1
Score (%)
Abstract Generalization
93.67Thinking Level · Extra High
--
94.00Thinking Level · High
ARC-AGI-2
Score (% solved); cost per task (USD)
Abstract Generalization
77.10Standard Mode
77.10Thinking Level · High
69.17Thinking Level · High
ARC-AGI-3 (Standard harness)
Action efficiency score(以 ARC Prize Standard harness 口径为准)
Abstract Generalization
0.21Thinking Level · High
0.42Thinking Level · High
0.51Thinking Level · High
HLE
Accuracy
Knowledge Exams
52.10Thinking Level · Extra High | Tools
51.40Thinking Level · High | Tools
53.00Extended Thinking | Tools
CritPt
Score
Scientific Reasoning
23.40Thinking Level · Extra High
17.70Thinking Enabled
12.60Thinking Level · High
GPQA Diamond
Accuracy
Scientific Reasoning
92.00Thinking Level · Extra High
94.30Thinking Level · High
91.31Extended Thinking
Creative Writing
Elo、大模型评判两两对战
Writing
1835.60Standard Mode
1488.90Standard Mode
1804.10Standard Mode
AIME 2026
Accuracy
Mathematics
99.17Thinking Level · Extra High | Tools
--
96.67Thinking Level · High | Tools
FrontierMath
Accuracy
Mathematics
47.60Thinking Level · Extra High
36.90Thinking Level · High
40.70Thinking Level · High
FrontierMath - Tier 4
Accuracy
Mathematics
27.10Thinking Level · Extra High
16.70Standard Mode
22.90Thinking Level · High
FrontierMath Tier 4 v2
Accuracy (verification_code)
Mathematics
49.00Thinking Level · Extra High
--
26.83Thinking Level · High
FrontierMath v2
Accuracy (verification_code)
Mathematics
78.60Thinking Level · Extra High
--
65.96Thinking Level · High
33 additional benchmarks remain in the chart above.

Standard API Pricing: GPT-5.4 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Gemini 3.1 Pro Preview: Base price applies to <= 200K
ModelSupplierStandard inputStandard outputBase price applies to
GPT-5.4
OpenAI$2.5 / 1M tokens$15 / 1M tokens—
Gemini 3.1 Pro Preview
Google DeepMind$2 / 1M tokens$12 / 1M tokens<= 200K
Claude Opus 4.6
Anthropic$5 / 1M tokens$25 / 1M tokens—

Version History

How each version of the GPT-5.4 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGPT-5.4CurrentGPT-5.2GPT-5.1
ARC-AGI-1
Score (%)
Abstract Generalization
93.67Thinking Level · Extra High
90.50Deep Thinking Mode
72.83Thinking Level · High
ARC-AGI-2
Score (% solved); cost per task (USD)
Abstract Generalization
77.10Standard Mode
54.20Deep Thinking Mode
17.64Thinking Level · High
HLE
Accuracy
Knowledge Exams
52.10Thinking Level · Extra High | Tools
45.50Deep Thinking Mode | Tools
42.70Thinking Level · High | Tools
CritPt
Score
Scientific Reasoning
23.40Thinking Level · Extra High
11.60Thinking Level · Extra High
4.90Thinking Level · High
GPQA Diamond
Accuracy
Scientific Reasoning
92.00Thinking Level · Extra High
93.20Deep Thinking Mode
88.10Thinking Enabled
Creative Writing
Elo、大模型评判两两对战
Writing
1835.60Standard Mode
1699.80Standard Mode
--
AIME 2026
Accuracy
Mathematics
99.17Thinking Level · Extra High | Tools
98.33Thinking Level · High | Tools
--
FrontierMath
Accuracy
Mathematics
47.60Thinking Level · Extra High
40.30Thinking Level · Extra High | Tools
26.70Thinking Level · High | Tools
FrontierMath - Tier 4
Accuracy
Mathematics
27.10Thinking Level · Extra High
18.80Thinking Level · Extra High
12.50Thinking Level · High | Tools
FrontierMath Tier 4 v2
Accuracy (verification_code)
Mathematics
49.00Thinking Level · Extra High
31.70Thinking Level · Extra High
--
FrontierMath v2
Accuracy (verification_code)
Mathematics
78.60Thinking Level · Extra High
67.40Thinking Level · Extra High
--
HMMT Feb 2026
Accuracy
Mathematics
97.73Thinking Level · Extra High | Tools
96.97Thinking Level · High | Tools
--
24 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: ARC-AGI-1 · Abstract Generalization

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the GPT-5.4 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
GPT-5.4
OpenAI$2.5 / 1M tokens$15 / 1M tokens—
GPT-5.2
OpenAI$1.75 / 1M tokens$14 / 1M tokens—
GPT-5.1
OpenAI$1.25 / 1M tokens$10 / 1M tokens—

Sources