DataLearner logo

Gemini 2.5 Pro Experimental 03-25 Benchmark Details

Gemini 2.5 Pro Experimental 03-25 currently shows benchmark results led by AIME 2024 (9 / 61, score 92), Aider-Polyglot (12 / 59, score 72.90), GeoBench ACW (5 / 20, score 81).

Benchmark Results

Gemini 2.5 Pro Experimental 03-25

Benchmark Results

Thinking
Tool usage
Internet

Knowledge Exams

2 evaluations
Benchmark / mode
Score
Rank/total
HLE
Standard Mode
18.80
162 / 233
HLE
unknown
18.16
165 / 233

Scientific Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
84
97 / 253

Honesty & Factuality

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleQA
Standard Mode
52.90
13 / 46

Repository Engineering

1 evaluations
Benchmark / mode
Score
Rank/total
SWE-bench Verified
Standard Mode
63.80
78 / 116

Mathematics

3 evaluations
Benchmark / mode
Score
Rank/total
AIME 2024
Standard Mode
92
9 / 61
AIME2025
Standard Mode
86.90
52 / 110
4.20
40 / 80

Algorithmic Coding

1 evaluations
Benchmark / mode
Score
Rank/total
LiveCodeBench
Standard Mode
70.40
57 / 125

Writing

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1393.40
65 / 106

Commonsense

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Standard Mode
51.60
52 / 96

Agentic Development

1 evaluations
Benchmark / mode
Score
Rank/total
Aider-Polyglot
Standard Mode
72.90
12 / 59

Tool Orchestration

2 evaluations
Benchmark / mode
Score
Rank/total
Claw Bench
Thinking EnabledTools
80.40
20 / 29
Pinch Bench
Thinking EnabledTools
71.90
30 / 38

Visual Understanding

2 evaluations
Benchmark / mode
Score
Rank/total
GeoBench ACW
Standard ModeToolsInternet
81
5 / 20
VPCT
Standard Mode
48
11 / 24

Long Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
Fiction.liveBench
Standard Mode
66.70
13 / 16

Logical Planning

1 evaluations
Benchmark / mode
Score
Rank/total
BALROG
Standard ModeTools
43.30
6 / 12