DataLearner logo

OpenAI o3-mini Benchmark Details

OpenAI o3-mini currently shows benchmark results led by MMLU (43 / 124, score 84.90), Aider-Polyglot (21 / 59, score 60.40), AIME2025 (53 / 110, score 86.50).

Benchmark Results

OpenAI o3-mini

Benchmark Results

Thinking
Tool usage

Knowledge Exams

2 evaluations
Benchmark / mode
Score
Rank/total
MMLU
Thinking Enabled
84.90
43 / 124
HLE
Thinking Enabled
13.40
184 / 235

Abstract Generalization

6 evaluations
Benchmark / mode
Score
Rank/total
ARC-AGI-1
Thinking Level · Low
14.50
180 / 193
ARC-AGI-1
Thinking Level · Medium
22.33
171 / 193
ARC-AGI-1
Thinking Level · High
34.50
160 / 193
ARC-AGI-2
Thinking Level · Low
0
176 / 181
ARC-AGI-2
Thinking Level · Medium
2.08
156 / 181
ARC-AGI-2
Thinking Level · High
3
152 / 181

Scientific Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Enabled
70.60
179 / 253
CritPt
Thinking Level · High
0.30
186 / 204

Repository Engineering

1 evaluations
Benchmark / mode
Score
Rank/total
SWE-bench Verified
Thinking Enabled
40.80
106 / 116

Mathematics

6 evaluations
Benchmark / mode
Score
Rank/total
MATH-500
Thinking Enabled
95.80
25 / 46
AIME2025
Thinking Enabled
86.50
53 / 110
AIME 2024
Thinking Enabled
60
42 / 61
FrontierMath v2
Thinking Level · High
18.60
52 / 58
FrontierMath - Tier 4
Thinking Level · High
4.20
40 / 80
FrontierMath Tier 4 v2
Thinking Level · High
0
41 / 42

Algorithmic Coding

1 evaluations
Benchmark / mode
Score
Rank/total
CodeForces
Thinking Enabled
2073
15 / 21

Commonsense

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Thinking Level · High
22.80
87 / 96

Agentic Development

4 evaluations
Benchmark / mode
Score
Rank/total
Aider-Polyglot
Thinking Level · Medium
53.80
29 / 59
Aider-Polyglot
Thinking Level · High
60.40
21 / 59
Terminal Bench Hard
Thinking EnabledTools
6.80
194 / 244
Terminal Bench Hard
Thinking Level · HighTools
6.10
204 / 244

Scientific Computing

1 evaluations
Benchmark / mode
Score
Rank/total
SciCode
Thinking Level · High
42.80
100 / 134

Service Workflows

1 evaluations
Benchmark / mode
Score
Rank/total
τ³-Banking
Thinking Level · HighTools
5.20
157 / 167

Documents & Charts

1 evaluations
Benchmark / mode
Score
Rank/total
GDP.pdf
Thinking Level · High
2
110 / 122

Capability Indices

1 evaluations
Benchmark / mode
Score
Rank/total
ECI
unknown
140.35
93 / 167