DataLearner logo

OpenAI o1-mini Benchmark Details

OpenAI o1-mini currently shows benchmark results led by HumanEval (8 / 140, score 92.40), MMLU (42 / 124, score 85.20), MMLU-Pro (64 / 176, score 80.30).

Benchmark Results

OpenAI o1-mini

Benchmark Results

Thinking

General Knowledge

3 evaluations
Benchmark / mode
Score
Rank/total
MMLU
Standard Mode
85.20
42 / 124
MMLU-Pro
Standard Mode
80.30
64 / 176
HLE
Thinking Enabled
3.60
547 / 565

Coding and Software Engineer

3 evaluations
Benchmark / mode
Score
Rank/total
HumanEval
Standard Mode
92.40
8 / 140
LiveCodeBench
Standard Mode
52
159 / 251
LiveCodeBench
Thinking Enabled
57.60
139 / 251

Other

2 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
60
381 / 463
GPQA Diamond
Thinking Enabled
60.30
378 / 463

Math and Reasoning

4 evaluations
Benchmark / mode
Score
Rank/total
MATH-500
Standard Mode
90
39 / 46
AIME 2024
Standard Mode
63.60
41 / 62
FrontierMath
Thinking Level · Medium
1.70
49 / 60
FrontierMath
Thinking Level · High
1.40
51 / 60

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Standard Mode
18.10
91 / 93

Agent Level Benchmark

1 evaluations
Benchmark / mode
Score
Rank/total
Aider-Polyglot
Standard Mode
32.90
42 / 59