DataLearner logo

OpenAI o3 Benchmark Details

OpenAI o3 currently shows benchmark results led by Aider-Polyglot (5 / 59, score 81.30), MMMU (9 / 74, score 82.90), MATH-500 (6 / 46, score 98.10).

Benchmark Results

OpenAI o3

Benchmark Results

Thinking
Tool usage

Knowledge Exams

4 evaluations
Benchmark / mode
Score
Rank/total
MMLU-Pro
Standard Mode
85.60
25 / 175
HLE
Thinking Level · Medium
19.20
161 / 235
HLE
Thinking Enabled
20.32
153 / 235
HLE
Thinking Level · High
20.32
153 / 235

Abstract Generalization

8 evaluations
Benchmark / mode
Score
Rank/total
ARC-AGI-1
Thinking Level · Low
41.50
151 / 193
ARC-AGI-1
Thinking Level · Medium
53.83
141 / 193
ARC-AGI-1
Thinking Enabled
60.80
126 / 193
ARC-AGI-1
Thinking Level · High
60.83
125 / 193
ARC-AGI-2
Thinking Level · Low
2
158 / 181
ARC-AGI-2
Thinking Level · Medium
2.98
153 / 181
ARC-AGI-2
Thinking Enabled
6.50
133 / 181
ARC-AGI-2
Thinking Level · High
6.53
132 / 181

Scientific Reasoning

4 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Level · Low
79.80
133 / 253
GPQA Diamond
Thinking Level · Medium
80.81
126 / 253
GPQA Diamond
Thinking Enabled
83.30
106 / 253
CritPt
Thinking Enabled
1.10
151 / 204

Honesty & Factuality

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleQA
Standard Mode
49.40
14 / 46

Repository Engineering

1 evaluations
Benchmark / mode
Score
Rank/total
SWE-bench Verified
Thinking Enabled
69.10
67 / 116

Mathematics

12 evaluations
Benchmark / mode
Score
Rank/total
MATH-500
Standard Mode
98.10
6 / 46
AIME 2024
Standard Mode
91.60
12 / 61
AIME2025
Thinking Enabled
88.90
46 / 110
FrontierMath v2
Thinking Level · Low
19.30
51 / 58
FrontierMath v2
Thinking Level · Medium
29.82
44 / 58
FrontierMath v2
Thinking Level · High
33.33
42 / 58
IMO-ProofBench
Thinking Enabled
20.50
11 / 16
IMO-ProofBench Advanced
Thinking Enabled
20.50
10 / 24
FrontierMath
Thinking Level · Low
10.30
25 / 60
FrontierMath
Thinking Level · Medium
10
28 / 60
FrontierMath
Thinking Level · High
10.30
25 / 60
FrontierMath - Tier 4
Thinking Level · High
2.10
56 / 80

Algorithmic Coding

2 evaluations
Benchmark / mode
Score
Rank/total
CodeClash
Standard ModeTools
1343
3 / 8
LiveCodeBench
Standard Mode
75.80
42 / 125

Writing

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1675.60
34 / 110

Agentic Development

4 evaluations
Benchmark / mode
Score
Rank/total
Aider-Polyglot
Standard Mode
76.90
9 / 59
Aider-Polyglot
Thinking Level · High
81.30
5 / 59
Terminal Bench Hard
Thinking EnabledTools
37.10
65 / 244
Terminal-Bench
Thinking Enabled
30.20
21 / 35

Visual Understanding

8 evaluations
Benchmark / mode
Score
Rank/total
MMMU
Standard Mode
82.90
9 / 74
MMMU
unknown
82.90
9 / 74
MMMU
Thinking Enabled
82.90
9 / 74
MMMU-Pro
unknown
76.40
73 / 229
MMMU-Pro
Thinking Enabled
70.10
130 / 229
GeoBench ACW
Thinking Level · Medium
74
11 / 20
GeoBench ACW
Thinking Level · High
60
19 / 20
VPCT
Thinking Level · Medium
52
10 / 24

Commonsense

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Thinking Level · High
53.10
50 / 96

Maintenance & Optimization

1 evaluations
Benchmark / mode
Score
Rank/total
GSO
Thinking Level · HighTools
8.80
15 / 21

ML Engineering

1 evaluations
Benchmark / mode
Score
Rank/total
WeirdML v2
Thinking Level · HighTools
52.42
36 / 52

Long Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
Fiction.liveBench
Thinking Level · Medium
88.90
6 / 16

Capability Frontier Metrics

1 evaluations
Benchmark / mode
Score
Rank/total
METR Time Horizons v1.1
Thinking Level · MediumTools
91.27
14 / 22

Clinical Workflows

2 evaluations
Benchmark / mode
Score
Rank/total
MedScribe
Thinking Level · HighTools
76.65
49 / 66
MedCode
Thinking Level · HighTools
47.29
25 / 64

Service Workflows

1 evaluations
Benchmark / mode
Score
Rank/total
SAGE
Thinking Level · HighTools
41.77
39 / 64

Capability Indices

1 evaluations
Benchmark / mode
Score
Rank/total
ECI
unknown
146.88
60 / 167