DataLearner logo

GPT-5.4 mini Benchmark Details

GPT-5.4 mini currently shows benchmark results led by Terminal Bench Hard (19 / 244, score 52.30), IF Bench (44 / 282, score 73.30), SAGE (10 / 64, score 50.81). This page also compares it with 2 competitor models and 1 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

GPT-5.4 mini

Benchmark Results

Thinking
Tool usage
Internet

Abstract Generalization

8 evaluations
Benchmark / mode
Score
Rank/total
ARC-AGI-1
LowTools
13
166 / 176
ARC-AGI-1
MediumTools
40.83
139 / 176
ARC-AGI-1
HighTools
58
120 / 176
ARC-AGI-1
Extra-HighTools
63.67
109 / 176
ARC-AGI-2
LowTools
1.11
155 / 164
ARC-AGI-2
MediumTools
4.44
133 / 164
ARC-AGI-2
HighTools
13.19
109 / 164
ARC-AGI-2
Extra-HighTools
18.90
102 / 164

Knowledge Exams

5 evaluations
Benchmark / mode
Score
Rank/total
HLE
Standard Mode
5.90
468 / 568
HLE
Medium
18.60
312 / 568
HLE
Extra-High
28.20
228 / 568
HLE
Extra-High
28.10
229 / 568
HLE
Extra-HighTools
41.50
123 / 568

Scientific Reasoning

5 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
64.14
363 / 461
82.30
204 / 461
GPQA Diamond
Extra-High
87.50
126 / 461
CritPt
Medium
2.90
119 / 204
CritPt
Extra-High
10
74 / 204

Writing

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1661.80
34 / 106

Mathematics

5 evaluations
Benchmark / mode
Score
Rank/total
FrontierMath v2
Standard Mode
17.19
54 / 58
24.56
47 / 58
FrontierMath v2
Extra-High
51.23
29 / 58
9.76
33 / 42
2.10
56 / 80

Repository Engineering

1 evaluations
Benchmark / mode
Score
Rank/total
SWE-Bench Pro - Public
Extra-HighTools
54.40
34 / 62

Service Workflows

5 evaluations
Benchmark / mode
Score
Rank/total
τ²-Bench - Telecom
Standard ModeTools
23.40
234 / 264
36.50
193 / 264
τ²-Bench - Telecom
Extra-HighTools
83.30
102 / 264
SAGE
Extra-HighTools
50.81
10 / 64
τ³-Banking
Extra-HighTools
25.60
75 / 167

Instruction Following

3 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Standard Mode
38.80
225 / 282
IF Bench
Medium
64.80
102 / 282
IF Bench
Extra-High
73.30
44 / 282

Cross-capability Suites

5 evaluations
Benchmark / mode
Score
Rank/total
LiveBench
Standard Mode
36.95
114 / 117
49.54
95 / 117
LiveBench
Medium
58.33
78 / 117
63.57
56 / 117
LiveBench
Deep Thinking Mode
66.37
50 / 117

Agentic Development

6 evaluations
Benchmark / mode
Score
Rank/total
Terminal Bench 2.0
Extra-HighTools
60
19 / 48
Terminal-Bench 2.1
Extra-HighTools
59.20
120 / 199
Terminal Bench Hard
Standard ModeTools
18.20
154 / 244
34.10
82 / 244
Terminal Bench Hard
Extra-HighTools
52.30
19 / 244
Terminal-Bench 4.0
Extra-HighTools
2
76 / 95

Tool Orchestration

4 evaluations
Benchmark / mode
Score
Rank/total
79.23
13 / 45
Claw Bench
Thinking EnabledTools
75.30
25 / 29
MCP-Atlas
Extra-HighTools
56.70
40 / 44
Tool Decathlon
Extra-HighTools
42.90
5 / 10

Memory & Persistence

5 evaluations
Benchmark / mode
Score
Rank/total
Context Arena
Standard Mode
23.34
119 / 126
39.26
99 / 126
47.64
91 / 126
49.60
90 / 126
Context Arena
Extra-High
50.92
88 / 126

Long Reasoning

3 evaluations
Benchmark / mode
Score
Rank/total
AA-LCR
Standard Mode
37
162 / 174
AA-LCR
Medium
67
120 / 174
AA-LCR
Extra-High
77
80 / 174

Desktop Workflows

1 evaluations
Benchmark / mode
Score
Rank/total
OSWorld-Verified
Extra-HighTools
72.10
20 / 28

Capability Indices

2 evaluations
Benchmark / mode
Score
Rank/total
ECI
unknown
148.83
53 / 167
Vals Index
Extra-HighTools
39.63
35 / 42

Cross-industry Work

2 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA v2
Standard ModeTools
734
106 / 110
GDPval-AA v2
Extra-HighTools
1095
86 / 110

Office & Business

1 evaluations
Benchmark / mode
Score
Rank/total
AA-Briefcase
Extra-HighTools
719
82 / 88

Visual Understanding

3 evaluations
Benchmark / mode
Score
Rank/total
MMMU-Pro
Standard Mode
60.50
174 / 229
MMMU-Pro
Medium
71.20
122 / 229
MMMU-Pro
Extra-High
73.30
110 / 229

Scientific Computing

1 evaluations
Benchmark / mode
Score
Rank/total
SciCode
Extra-High
52.10
60 / 134

Legal

3 evaluations
Benchmark / mode
Score
Rank/total
Harvey Lab-AA
Extra-HighTools
60.75
37 / 44
Legal Research Bench
Extra-HighTools
12.50
38 / 42
0
36 / 43

Finance

2 evaluations
Benchmark / mode
Score
Rank/total
EMB (Excel Modeling)
Extra-HighTools
45.43
35 / 42
Finance Agent v2
Extra-HighTools
45.36
35 / 42

Documents & Charts

1 evaluations
Benchmark / mode
Score
Rank/total
GDP.pdf
Extra-High
13.40
66 / 122

Data Analysis

1 evaluations
Benchmark / mode
Score
Rank/total
AA-AnalystAgent
Extra-HighToolsInternet
15
21 / 29

Maintenance & Optimization

1 evaluations
Benchmark / mode
Score
Rank/total
Code Migration
Extra-HighTools
12.94
37 / 43

Code Generation & Editing

1 evaluations
Benchmark / mode
Score
Rank/total
Vibe Code Bench v1.1
Extra-HighTools
47.97
38 / 60

Competitor Comparison

Benchmark scores for GPT-5.4 mini compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGPT-5.4 miniCurrentHaiku 4.5Gemini 3.0 Flash
ARC-AGI-1
Score (%)
Abstract Generalization
63.67Thinking Level · Extra High | Tools
47.67Extended Thinking
--
ARC-AGI-2
Score (% solved); cost per task (USD)
Abstract Generalization
18.90Thinking Level · Extra High | Tools
4.50Extended Thinking
33.60Thinking Enabled
HLE
Accuracy
Knowledge Exams
41.50Thinking Level · Extra High | Tools
10.40Thinking Enabled
43.50Thinking Enabled | Tools
CritPt
Score
Scientific Reasoning
10.00Thinking Level · Extra High
--
8.60Thinking Enabled
GPQA Diamond
Accuracy
Scientific Reasoning
87.50Thinking Level · Extra High
73.30Extended Thinking
90.40Thinking Enabled
FrontierMath - Tier 4
Accuracy
Mathematics
2.10Thinking Level · High
2.1032K
4.20Standard Mode
SWE-Bench Pro - Public
Accuracy
Repository Engineering
54.40Thinking Level · Extra High | Tools
39.45Extended Thinking | Tools
49.60Thinking Level · High | Tools
SAGE
Accuracy (%)
Service Workflows
50.81Thinking Level · Extra High | Tools
31.82Thinking Enabled | Tools
--
τ²-Bench - Telecom
Accuracy
Service Workflows
83.30Thinking Level · Extra High | Tools
54.70Thinking Enabled | Tools
91.23Thinking Level · High | Tools
τ³-Banking
Score
Service Workflows
25.60Thinking Level · Extra High | Tools
9.30Thinking Enabled | Tools
27.32Thinking Level · High | Tools
IF Bench
Accuracy
Instruction Following
73.30Thinking Level · Extra High
54.30Thinking Enabled
78.00Thinking Enabled
LiveBench
Accuracy
Cross-capability Suites
66.37Deep Thinking Mode
61.3264K
72.40Thinking Level · High
21 additional benchmarks remain in the chart above.

Standard API Pricing: GPT-5.4 mini vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
GPT-5.4 mini
OpenAI$0.75 / 1M tokens$4.5 / 1M tokens—
Haiku 4.5
Anthropic$1 / 1M tokens$5 / 1M tokens—
Gemini 3.0 Flash
Google DeepMind$0.5 / 1M tokens$3 / 1M tokens—

Version History

How each version of the GPT-5.4 mini series stacks up on benchmark tests

GPT-5.4 miniGPT-5-mini
Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGPT-5.4 miniCurrentGPT-5-mini
ARC-AGI-1
Score (%)
Abstract Generalization
63.67Thinking Level · Extra High | Tools
54.33Thinking Level · High | Tools
ARC-AGI-2
Score (% solved); cost per task (USD)
Abstract Generalization
18.90Thinking Level · Extra High | Tools
4.44Thinking Level · High | Tools
HLE
Accuracy
Knowledge Exams
41.50Thinking Level · Extra High | Tools
21.50Thinking Level · High
CritPt
Score
Scientific Reasoning
10.00Thinking Level · Extra High
1.40Thinking Level · Medium
GPQA Diamond
Accuracy
Scientific Reasoning
87.50Thinking Level · Extra High
82.80Thinking Level · High
Creative Writing
Elo、大模型评判两两对战
Writing
1661.80Standard Mode
1310.30Standard Mode
FrontierMath - Tier 4
Accuracy
Mathematics
2.10Thinking Level · High
6.30Thinking Level · High
FrontierMath Tier 4 v2
Accuracy (verification_code)
Mathematics
9.76Thinking Level · Extra High
12.20Thinking Level · High
FrontierMath v2
Accuracy (verification_code)
Mathematics
51.23Thinking Level · Extra High
46.67Thinking Level · High
SAGE
Accuracy (%)
Service Workflows
50.81Thinking Level · Extra High | Tools
42.99Thinking Level · High | Tools
τ²-Bench - Telecom
Accuracy
Service Workflows
83.30Thinking Level · Extra High | Tools
71.10Thinking Level · Medium | Tools
τ³-Banking
Score
Service Workflows
25.60Thinking Level · Extra High | Tools
15.50Thinking Level · High | Tools
12 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: ARC-AGI-1 · Abstract Generalization

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the GPT-5.4 mini Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
GPT-5.4 mini
OpenAI$0.75 / 1M tokens$4.5 / 1M tokens—
GPT-5-mini
OpenAI$0.25 / 1M tokens$2 / 1M tokens—