DataLearner logo

GPT-4o(2025-03-27) Benchmark Details

GPT-4o(2025-03-27) currently shows benchmark results led by MMLU Pro (66 / 176, score 79.80), SimpleQA (20 / 47, score 40.30), Creative Writing (50 / 106, score 1502.10).

Benchmark Results

GPT-4o(2025-03-27)

Benchmark Results

Thinking

General Knowledge

2 evaluations
Benchmark / mode
Score
Rank/total
MMLU Pro
Standard Mode
79.80
66 / 176
ARC-AGI-1
Standard Mode
8.80
141 / 147

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
66.90
346 / 462

Common Sense

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleQA
Standard Mode
40.30
20 / 47

Coding and Software Engineer

1 evaluations
Benchmark / mode
Score
Rank/total
LiveCodeBench
Standard Mode
35.80
195 / 250

Math and Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
AIME2025
Standard Mode
26.70
189 / 215

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1502.10
50 / 106

Agent Level Benchmark

1 evaluations
Benchmark / mode
Score
Rank/total
Aider-Polyglot
Standard Mode
45.30
36 / 59