DataLearner logo

OpenAI o1 Benchmark Details

OpenAI o1 currently shows benchmark results led by MMLU Pro (1 / 133, score 91.04), MMLU (4 / 124, score 91.80), MATH (2 / 42, score 96.40). 1 source link is attached for reference.

Benchmark Results

OpenAI o1

Benchmark Results

Thinking
Tool usage

General Knowledge

3 evaluations
Benchmark / mode
Score
Rank/total
MMLU
Standard Mode
91.80
4 / 124
MMLU Pro
Standard Mode
91.04
1 / 133
HLE
Standard Mode
9.10
163 / 185

Math and Reasoning

4 evaluations
Benchmark / mode
Score
Rank/total
MATH
Standard Mode
96.40
2 / 42
MATH-500
Standard Mode
96.40
17 / 44
AIME 2024
Standard Mode
79.20
31 / 62
FrontierMath
Thinking Level · High
9.30
30 / 60

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
77.30
159 / 270

Common Sense

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleQA
Standard Mode
42.60
19 / 47

Coding and Software Engineer

4 evaluations
Benchmark / mode
Score
Rank/total
LiveCodeBench
Standard Mode
71
51 / 127
SWE-bench Verified
Standard Mode
48.90
100 / 114
SWE-bench Verified
Thinking Level · High
41
103 / 114
WeirdML v2
Thinking Level · HighTools
43.82
43 / 52

Common Sense Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Thinking Level · Medium
36.70
47 / 67
SimpleBench
Thinking Level · High
40.10
43 / 67

Agent Level Benchmark

2 evaluations
Benchmark / mode
Score
Rank/total
Aider-Polyglot
Thinking Level · High
61.70
19 / 59
METR Time Horizons v1.1
Thinking Level · MediumTools
39.21
18 / 22

Multimodal Understanding

2 evaluations
Benchmark / mode
Score
Rank/total
GeoBench ACW
Thinking Level · Medium
80
8 / 20
VPCT
Thinking Level · Medium
37
21 / 24

Other

1 evaluations
Benchmark / mode
Score
Rank/total
Fiction.liveBench
Thinking Level · Medium
83.30
8 / 16

Sources