DataLearner logo

OpenAI o4 - mini Benchmark Details

OpenAI o4 - mini currently shows benchmark results led by AIME 2024 (1 / 62, score 98.70), MMLU (2 / 124, score 93), AIME2025 (11 / 216, score 99.50).

Benchmark Results

OpenAI o4 - mini

Benchmark Results

Thinking
Tool usage

General Knowledge

9 evaluations
Benchmark / mode
Score
Rank/total
MMLU
Thinking Enabled
93
2 / 124
MMLU-Pro
Thinking Enabled
80.60
60 / 176
ARC-AGI-1
Thinking Enabled
58.70
97 / 147
HLE
Thinking Level · Medium
14.28
345 / 565
HLE
Thinking Enabled
14.28
345 / 565
HLE
Thinking EnabledTools
17.70
316 / 565
HLE
Thinking Level · High
18.08
314 / 565
HLE
Thinking Level · High
16.50
326 / 565
CritPt
Thinking Level · High
0.60
169 / 201

Other

4 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Level · Low
75.25
278 / 463
GPQA Diamond
Thinking Level · Medium
77.78
259 / 463
GPQA Diamond
Thinking Enabled
81.40
213 / 463
GPQA Diamond
Thinking Level · High
78.40
254 / 463

Coding and Software Engineer

5 evaluations
Benchmark / mode
Score
Rank/total
CodeForces
Thinking EnabledTools
2719
7 / 21
LiveCodeBench
Thinking Level · High
85.90
26 / 251
SWE-bench Verified
Thinking Enabled
68.10
70 / 116
WeirdML v2
Thinking Level · HighTools
52.56
35 / 52
GSO
Thinking Level · HighTools
3.60
20 / 21

Math and Reasoning

18 evaluations
Benchmark / mode
Score
Rank/total
AIME2025
Thinking Enabled
92.70
42 / 216
AIME2025
Thinking EnabledTools
99.50
11 / 216
AIME2025
Thinking Level · High
90.70
52 / 216
AIME 2024
Thinking Enabled
93.40
5 / 62
AIME 2024
Thinking EnabledTools
98.70
1 / 62
FrontierMath v2
Thinking Level · Low
16.14
55 / 58
FrontierMath v2
Thinking Level · Medium
28.77
45 / 58
FrontierMath v2
Thinking Level · High
36.14
38 / 58
FrontierMath
Thinking Level · Low
9.70
29 / 60
FrontierMath
Thinking Level · Medium
19.30
18 / 60
FrontierMath
Thinking Level · High
17.20
21 / 60
IMO-ProofBench
Thinking Level · High
11.40
12 / 16
IMO-ProofBench Advanced
Thinking Level · High
11.40
15 / 24
IMO 2024
Thinking Enabled
7.70
7 / 10
FrontierMath - Tier 4
Thinking Level · Medium
2.10
56 / 80
FrontierMath - Tier 4
Thinking Level · High
6.30
35 / 80
FrontierMath Tier 4 v2
Thinking Level · High
4.88
34 / 42
IMO 2025
Thinking Enabled
3
7 / 9

Multimodal Understanding

5 evaluations
Benchmark / mode
Score
Rank/total
MMMU
unknown
81.60
14 / 74
MMMU-Pro
Thinking Level · High
69.20
134 / 229
GeoBench ACW
Thinking Level · Medium
64
15 / 20
GeoBench ACW
Thinking Level · High
64
15 / 20
VPCT
Thinking Level · Medium
57.50
8 / 24

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Thinking Level · High
38.70
69 / 93

Agent Level Benchmark

6 evaluations
Benchmark / mode
Score
Rank/total
METR Time Horizons v1.1
Thinking Level · MediumTools
76.51
15 / 22
Aider-Polyglot
Thinking Level · High
72
13 / 59
τ²-Bench
Thinking EnabledTools
56.90
32 / 44
τ²-Bench - Telecom
Thinking EnabledTools
50.20
171 / 264
τ²-Bench - Telecom
Thinking Level · HighTools
55.60
162 / 264
Terminal Bench Hard
Thinking Level · HighTools
15.20
168 / 244

Instruction Following

1 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Thinking Level · High
68.70
82 / 282

Other

1 evaluations
Benchmark / mode
Score
Rank/total
Fiction.liveBench
Thinking Level · Medium
77.80
9 / 16