DataLearner logo

GPT-4o Benchmark Details

GPT-4o currently shows benchmark results led by HumanEval (8 / 101, score 90), MMLU (16 / 124, score 88.70), BBH (5 / 21, score 91.70).

Benchmark Results

GPT-4o

Benchmark Results

Thinking
Tool usage

General Knowledge

4 evaluations
Benchmark / mode
Score
Rank/total
BBH
Standard Mode
91.70
5 / 21
MMLU
Standard Mode
88.70
16 / 124
MMLU Pro
Standard Mode
77.90
76 / 133
HLE
Standard Mode
5.30
177 / 185

Coding and Software Engineer

4 evaluations
Benchmark / mode
Score
Rank/total
HumanEval
Standard Mode
90
8 / 101
LiveCodeBench
Standard Mode
35.10
112 / 127
SWE-bench Verified
Standard Mode
31
109 / 114
23.30
6 / 8

Math and Reasoning

5 evaluations
Benchmark / mode
Score
Rank/total
MATH
Standard Mode
75.90
16 / 42
MATH-500
Standard Mode
75.90
43 / 44
AIME2025
Standard ModeTools
42.10
93 / 106
AIME 2024
Standard Mode
9.30
61 / 62
FrontierMath
Standard Mode
0.30
57 / 60

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
70.10
195 / 270

Common Sense

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleQA
Standard Mode
38.20
22 / 47

Agent Level Benchmark

1 evaluations
Benchmark / mode
Score
Rank/total
Aider-Polyglot
Standard Mode
23.10
47 / 59

Claw-style Agent Evaluation

1 evaluations
Benchmark / mode
Score
Rank/total
Pinch Bench
Thinking ModeTools
71.10
31 / 38