DataLearner logo

o3-pro Benchmark Details

o3-pro currently shows benchmark results led by Aider-Polyglot (3 / 59, score 84.90), Fiction.liveBench (1 / 16, score 97.20), AIME 2024 (8 / 62, score 93).

Benchmark Results

o3-pro

Benchmark Results

Thinking
Tool usage

General Knowledge

2 evaluations
Benchmark / mode
Score
Rank/total
ARC-AGI
Standard Mode
59.30
38 / 68
HLE
Standard Mode
21
128 / 185

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
84
108 / 270

Coding and Software Engineer

2 evaluations
Benchmark / mode
Score
Rank/total
SWE-bench Verified
Thinking Level · High
75
39 / 114
WeirdML v2
Thinking Level · HighTools
58.21
30 / 52

Math and Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
AIME 2024
Standard Mode
93
8 / 62
AIME2025
Standard Mode
93
30 / 106

Agent Level Benchmark

1 evaluations
Benchmark / mode
Score
Rank/total
Aider-Polyglot
Thinking Level · High
84.90
3 / 59

AI Agent - Tool Usage

1 evaluations
Benchmark / mode
Score
Rank/total
MCP-Atlas
Standard ModeTools
44.50
39 / 40

Other

1 evaluations
Benchmark / mode
Score
Rank/total
Fiction.liveBench
Thinking Level · Medium
97.20
1 / 16