DataLearner logo

GPT-5-Nano Benchmark Details

GPT-5-Nano currently shows benchmark results led by FrontierMath (31 / 60, score 8.30), AIME2025 (57 / 110, score 85), ECI (98 / 167, score 139.39).

Benchmark Results

GPT-5-Nano

Benchmark Results

Thinking
Tool usage

Abstract Generalization

6 evaluations
Benchmark / mode
Score
Rank/total
ARC-AGI-1
Thinking Level · LowTools
4.04
192 / 193
ARC-AGI-1
Thinking Level · MediumTools
20.71
174 / 193
ARC-AGI-1
Thinking Level · HighTools
16.67
177 / 193
ARC-AGI-2
Thinking Level · LowTools
0
176 / 181
ARC-AGI-2
Thinking Level · MediumTools
0.88
172 / 181
ARC-AGI-2
Thinking Level · HighTools
2.61
154 / 181

Scientific Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Level · Low
57.58
217 / 253

Mathematics

9 evaluations
Benchmark / mode
Score
Rank/total
AIME2025
Thinking Level · LowTools
47.22
98 / 110
AIME2025
Thinking Level · HighTools
85
57 / 110
FrontierMath v2
Thinking Level · Low
5.96
57 / 58
FrontierMath v2
Thinking Level · High
20
50 / 58
FrontierMath
Thinking Level · Medium
7.20
33 / 60
FrontierMath
Thinking Level · High
8.30
31 / 60
FrontierMath Tier 4 v2
Thinking Level · High
2.44
37 / 42
FrontierMath - Tier 4
Thinking Level · Medium
2.10
56 / 80
FrontierMath - Tier 4
Thinking Level · High
0
72 / 80

Writing

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
704.90
107 / 110

Visual Understanding

4 evaluations
Benchmark / mode
Score
Rank/total
MMMU-Pro
Standard Mode
31.80
226 / 229
MMMU-Pro
Thinking Level · Medium
58.20
183 / 229
MMMU-Pro
Thinking Level · High
61
173 / 229
MMMU
Standard Mode
57.60
63 / 74

Documents & Charts

1 evaluations
Benchmark / mode
Score
Rank/total
DocVQA
Standard Mode
78.30
5 / 5

Cross-capability Suites

3 evaluations
Benchmark / mode
Score
Rank/total
LiveBench
Standard Mode
48.56
100 / 117
LiveBench
Thinking Level · Low
34.34
115 / 117
LiveBench
Thinking Level · High
48.62
99 / 117

Agentic Development

3 evaluations
Benchmark / mode
Score
Rank/total
Terminal Bench Hard
Standard ModeTools
6.80
194 / 244
Terminal Bench Hard
Thinking Level · MediumTools
17.40
158 / 244
Terminal Bench Hard
Thinking Level · HighTools
12.10
178 / 244

Tool Orchestration

1 evaluations
Benchmark / mode
Score
Rank/total
Pinch Bench
Thinking ModeTools
68.80
34 / 38

Clinical Workflows

2 evaluations
Benchmark / mode
Score
Rank/total
MedScribe
Thinking Level · HighTools
72.86
56 / 66
MedCode
Thinking Level · HighTools
30.44
59 / 64

Service Workflows

1 evaluations
Benchmark / mode
Score
Rank/total
SAGE
Thinking Level · HighTools
30.38
59 / 64

Capability Indices

1 evaluations
Benchmark / mode
Score
Rank/total
ECI
unknown
139.39
98 / 167