DataLearner logo

GPT-5.4 nano Benchmark Details

GPT-5.4 nano currently shows benchmark results led by Terminal Bench Hard (45 / 244, score 42.40), LiveBench (39 / 117, score 69.58), Claw Bench (10 / 29, score 89.70).

Benchmark Results

GPT-5.4 nano

Benchmark Results

Thinking
Tool usage

Abstract Generalization

8 evaluations
Benchmark / mode
Score
Rank/total
ARC-AGI-1
LowTools
18.33
175 / 193
ARC-AGI-1
MediumTools
33
164 / 193
ARC-AGI-1
HighTools
38.17
155 / 193
ARC-AGI-1
Extra-HighTools
51.50
144 / 193
ARC-AGI-2
LowTools
1.53
165 / 181
ARC-AGI-2
MediumTools
1.94
159 / 181
ARC-AGI-2
HighTools
3.61
151 / 181
ARC-AGI-2
Extra-HighTools
5.69
137 / 181

Knowledge Exams

2 evaluations
Benchmark / mode
Score
Rank/total
HLE
Extra-High
24.30
141 / 235
HLE
Extra-HighTools
37.70
88 / 235

Scientific Reasoning

4 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
55.56
221 / 253
72.22
169 / 253
CritPt
Medium
5.10
98 / 204
CritPt
Extra-High
9.30
79 / 204

Visual Understanding

5 evaluations
Benchmark / mode
Score
Rank/total
MMMU
Extra-High
66.10
53 / 74
MMMU
Extra-HighTools
69.50
47 / 74
MMMU-Pro
Standard Mode
43.80
217 / 229
MMMU-Pro
Medium
59.50
179 / 229
MMMU-Pro
Extra-High
65.40
152 / 229

Mathematics

5 evaluations
Benchmark / mode
Score
Rank/total
FrontierMath v2
Standard Mode
4.56
58 / 58
20.35
49 / 58
44.91
32 / 58
12.20
31 / 42
6.30
35 / 80

Repository Engineering

1 evaluations
Benchmark / mode
Score
Rank/total
SWE-Bench Pro - Public
Extra-HighTools
52.40
41 / 62

Cross-capability Suites

5 evaluations
Benchmark / mode
Score
Rank/total
LiveBench
Standard Mode
32.39
117 / 117
48.67
98 / 117
LiveBench
Medium
58.46
77 / 117
62.75
58 / 117
LiveBench
Deep Thinking Mode
69.58
39 / 117

Agentic Development

5 evaluations
Benchmark / mode
Score
Rank/total
Terminal Bench 2.0
Extra-HighTools
46.30
42 / 48
Terminal Bench Hard
Standard ModeTools
24.20
134 / 244
33.30
86 / 244
Terminal Bench Hard
Extra-HighTools
42.40
45 / 244
Terminal-Bench 4.0
Extra-HighTools
0.50
93 / 98

Tool Orchestration

3 evaluations
Benchmark / mode
Score
Rank/total
Claw Bench
Thinking EnabledTools
89.70
10 / 29
68.99
27 / 45
Tool Decathlon
Extra-HighTools
35.50
9 / 10

Memory & Persistence

5 evaluations
Benchmark / mode
Score
Rank/total
Context Arena
Standard Mode
13.04
126 / 126
23.31
120 / 126
32.74
113 / 126
39.03
100 / 126
Context Arena
Extra-High
52.79
85 / 126

Desktop Workflows

1 evaluations
Benchmark / mode
Score
Rank/total
OSWorld-Verified
Extra-HighTools
39
27 / 28

Capability Indices

2 evaluations
Benchmark / mode
Score
Rank/total
ECI
unknown
145.78
73 / 167
Vals Index
HighTools
33
37 / 42

Scientific Computing

1 evaluations
Benchmark / mode
Score
Rank/total
SciCode
Extra-High
47.20
83 / 134

Service Workflows

2 evaluations
Benchmark / mode
Score
Rank/total
SAGE
HighTools
38.08
45 / 64
τ³-Banking
Extra-HighTools
27.40
69 / 167

Legal

3 evaluations
Benchmark / mode
Score
Rank/total
Harvey Lab-AA
Extra-HighTools
52.24
39 / 44
6.25
41 / 42

Finance

2 evaluations
Benchmark / mode
Score
Rank/total
44.75
36 / 42
38.22
38 / 43

Documents & Charts

1 evaluations
Benchmark / mode
Score
Rank/total
GDP.pdf
Extra-High
7.80
92 / 122

Maintenance & Optimization

1 evaluations
Benchmark / mode
Score
Rank/total
14.47
35 / 43

Code Generation & Editing

1 evaluations
Benchmark / mode
Score
Rank/total
26.10
44 / 61

Clinical Workflows

2 evaluations
Benchmark / mode
Score
Rank/total
MedScribe
HighTools
77.09
48 / 66
MedCode
HighTools
41.03
43 / 64