DataLearner logo

OpenAI o3 Benchmark Details

OpenAI o3 currently shows benchmark results led by Aider-Polyglot (5 / 59, score 81.30), Creative Writing (2 / 23, score 87.65), MATH-500 (5 / 44, score 98.10).

Benchmark Results

OpenAI o3

Benchmark Results

Thinking
Tool usage

General Knowledge

5 evaluations
Benchmark / mode
Score
Rank/total
85.60
23 / 132
83.30
64 / 187
60.80
37 / 68
20.32
116 / 172
6.50
44 / 62

Common Sense

1 evaluations
Benchmark / mode
Score
Rank/total
49.40
14 / 47

Coding and Software Engineer

3 evaluations
Benchmark / mode
Score
Rank/total
CodeClash
Standard ModeTools
1343
3 / 8
75.80
39 / 123
69.10
65 / 112

Math and Reasoning

8 evaluations
Benchmark / mode
Score
Rank/total
98.10
5 / 44
91.60
12 / 62
88.90
42 / 107
20.50
11 / 16
10.30
25 / 60
10.30
25 / 60
10
28 / 60
FrontierMath - Tier 4
Thinking Level · High
2.10
56 / 80

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
87.65
2 / 23

AI Agent - Tool Usage

1 evaluations
Benchmark / mode
Score
Rank/total
30.20
21 / 35

Multimodal Understanding

2 evaluations
Benchmark / mode
Score
Rank/total
82.90
7 / 29
82.90
7 / 29

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
Simple Bench
Thinking Level · High
53.10
25 / 63

Agent Level Benchmark

2 evaluations
Benchmark / mode
Score
Rank/total
Aider-Polyglot
Standard Mode
76.90
9 / 59
Aider-Polyglot
Thinking Level · High
81.30
5 / 59