DataLearner logo

GPT-5-mini Benchmark Details

GPT-5-mini currently shows benchmark results led by FrontierMath (18 / 60, score 19.30), Terminal Bench Hard (86 / 244, score 33.30), MMLU-Pro (74 / 175, score 78).

Benchmark Results

GPT-5-mini

Benchmark Results

Thinking
Tool usage

Knowledge Exams

3 evaluations
Benchmark / mode
Score
Rank/total
MMLU-Pro
Thinking Enabled
78
74 / 175
HLE
unknown
19.44
159 / 235
HLE
Thinking Enabled
5
227 / 235

Abstract Generalization

6 evaluations
Benchmark / mode
Score
Rank/total
ARC-AGI-1
Thinking Level · LowTools
26.33
168 / 193
ARC-AGI-1
Thinking Level · MediumTools
37.33
157 / 193
ARC-AGI-1
Thinking Level · HighTools
54.33
140 / 193
ARC-AGI-2
Thinking Level · LowTools
0.83
173 / 181
ARC-AGI-2
Thinking Level · MediumTools
4.03
148 / 181
ARC-AGI-2
Thinking Level · HighTools
4.44
146 / 181

Scientific Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Enabled
69
187 / 253
CritPt
Thinking Level · Medium
1.40
139 / 204

Algorithmic Coding

2 evaluations
Benchmark / mode
Score
Rank/total
CodeClash
Standard ModeTools
1200
5 / 8
LiveCodeBench
Thinking Enabled
55
88 / 125

Mathematics

11 evaluations
Benchmark / mode
Score
Rank/total
AIME2025
Standard Mode
47
99 / 110
AIME2025
Thinking Enabled
47
99 / 110
AIME2025
Thinking Level · HighTools
87.50
49 / 110
FrontierMath v2
Thinking Level · Low
18.25
53 / 58
FrontierMath v2
Thinking Level · High
46.67
30 / 58
FrontierMath
Thinking Level · Medium
19.30
18 / 60
FrontierMath
Thinking Level · High
19
20 / 60
FrontierMath Tier 4 v2
Thinking Level · High
12.20
31 / 42
FrontierMath - Tier 4
Thinking Level · Medium
4.20
40 / 80
FrontierMath - Tier 4
Thinking Level · High
6.30
35 / 80
MathArena Apex
Thinking Level · HighTools
1.04
14 / 17

Writing

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1313
76 / 110

Agentic Development

4 evaluations
Benchmark / mode
Score
Rank/total
Terminal Bench Hard
Standard ModeTools
14.40
170 / 244
Terminal Bench Hard
Thinking Level · MediumTools
28.80
116 / 244
Terminal Bench Hard
Thinking Level · HighTools
33.30
86 / 244
Terminal-Bench
Thinking Enabled
14
33 / 35

Cross-capability Suites

3 evaluations
Benchmark / mode
Score
Rank/total
LiveBench
Standard Mode
61
68 / 117
LiveBench
Thinking Level · Low
53.07
87 / 117
LiveBench
Thinking Level · High
65.91
51 / 117

Tool Orchestration

1 evaluations
Benchmark / mode
Score
Rank/total
Pinch Bench
Thinking EnabledTools
80.30
24 / 38

Visual Understanding

5 evaluations
Benchmark / mode
Score
Rank/total
MMMU-Pro
Standard Mode
58.40
182 / 229
MMMU-Pro
Thinking Level · Medium
68.80
141 / 229
MMMU-Pro
Thinking Level · High
70.10
130 / 229
VPCT
Thinking Level · Medium
40.20
14 / 24
VPCT
Thinking Level · High
39
18 / 24

Scientific Computing

1 evaluations
Benchmark / mode
Score
Rank/total
SciCode
Thinking Level · High
39
110 / 134

Service Workflows

2 evaluations
Benchmark / mode
Score
Rank/total
SAGE
Thinking Level · HighTools
43
38 / 64
τ³-Banking
Thinking Level · HighTools
15.50
109 / 167

ML Engineering

1 evaluations
Benchmark / mode
Score
Rank/total
WeirdML v2
Thinking Level · HighTools
52.67
34 / 52

Long Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
Fiction.liveBench
Thinking Level · Medium
69.40
11 / 16

Documents & Charts

1 evaluations
Benchmark / mode
Score
Rank/total
GDP.pdf
Thinking Level · High
8.40
90 / 122

Code Generation & Editing

1 evaluations
Benchmark / mode
Score
Rank/total
Vibe Code Bench v1.1
Thinking Level · HighTools
14.17
53 / 61

Clinical Workflows

2 evaluations
Benchmark / mode
Score
Rank/total
MedScribe
Thinking Level · HighTools
80.58
39 / 66
MedCode
Thinking Level · HighTools
43.05
36 / 64

Capability Indices

1 evaluations
Benchmark / mode
Score
Rank/total
ECI
unknown
145.51
75 / 167