DataLearner logo

Gemini 2.5 Flash Benchmark Details

Gemini 2.5 Flash currently shows benchmark results led by AIME 2024 (16 / 62, score 88), GPQA Diamond (196 / 463, score 82.80), GeoBench ACW (9 / 20, score 76).

Benchmark Results

Gemini 2.5 Flash

Benchmark Results

Thinking
Tool usage

General Knowledge

9 evaluations
Benchmark / mode
Score
Rank/total
LiveBench
Thinking Level · High
47.74
103 / 117
ARC-AGI-1
Standard Mode
32.30
125 / 147
HLE
Standard Mode
8.40
415 / 565
HLE
Standard Mode
4.70
492 / 565
HLE
unknown
12.08
369 / 565
HLE
Thinking Enabled
12.10
367 / 565
HLE
Thinking Enabled
11
379 / 565
CritPt
Standard Mode
1.40
136 / 201
CritPt
Thinking Enabled
1.10
148 / 201

Other

2 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
78.30
255 / 463
GPQA Diamond
Thinking Enabled
82.80
196 / 463

Common Sense

2 evaluations
Benchmark / mode
Score
Rank/total
SimpleQA
Standard Mode
25.80
30 / 47
SimpleQA
Thinking Enabled
26.90
29 / 47

Coding and Software Engineer

5 evaluations
Benchmark / mode
Score
Rank/total
LiveCodeBench
Standard Mode
41.10
183 / 251
LiveCodeBench
Thinking Enabled
55.40
151 / 251
SWE-bench Verified
Standard Mode
50
96 / 116
SWE-bench Verified
Thinking Enabled
48.90
100 / 116
WeirdML v2
16KTools
40.95
46 / 52

Math and Reasoning

5 evaluations
Benchmark / mode
Score
Rank/total
AIME 2024
Standard Mode
88
16 / 62
AIME2025
Standard Mode
61.60
137 / 216
AIME2025
Thinking Enabled
72
117 / 216
IMO 2024
Standard Mode
7.80
6 / 10
4.20
40 / 80

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1134.60
84 / 106

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Standard Mode
41.20
64 / 93

Agent Level Benchmark

7 evaluations
Benchmark / mode
Score
Rank/total
Aider-Polyglot
Standard Mode
47.10
35 / 59
55.10
27 / 59
BALROG
Standard ModeTools
33.50
9 / 12
τ²-Bench - Telecom
Standard ModeTools
14.90
258 / 264
τ²-Bench - Telecom
Thinking EnabledTools
31.60
207 / 264
Terminal Bench Hard
Standard ModeTools
12.10
178 / 244
Terminal Bench Hard
Thinking EnabledTools
13.60
173 / 244

Instruction Following

2 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Standard Mode
39
223 / 282
IF Bench
Thinking Enabled
50.30
156 / 282

Long Context

1 evaluations
Benchmark / mode
Score
Rank/total
AA-LCR
Standard Mode
49.90
139 / 171

Claw-style Agent Evaluation

1 evaluations
Benchmark / mode
Score
Rank/total
Pinch Bench
Thinking EnabledTools
70.70
32 / 38

Multimodal Understanding

3 evaluations
Benchmark / mode
Score
Rank/total
GeoBench ACW
Standard Mode
76
9 / 20
MMMU-Pro
Standard Mode
65.50
151 / 229
MMMU-Pro
Thinking Enabled
69.10
138 / 229

Other

1 evaluations
Benchmark / mode
Score
Rank/total
Fiction.liveBench
Standard Mode
77.80
9 / 16