DataLearner logo

Gemini 3.0 FlashvsGemini 2.5 Flash

Across 12 shared benchmarks, Gemini 3.0 Flash leads overall: Gemini 3.0 Flash wins 11, Gemini 2.5 Flash wins 0, with 1 ties and an average score difference of +18.21.

Google Deep Mind
Gemini 3.0 Flash

Google Deep Mind · 2025-12-17 · Chat model

Google Deep Mind
Gemini 2.5 Flash

Google Deep Mind · 2025-04-17 · Reasoning model

Gemini 3.0 Flash11 wins(92%)Ties1(0%)0 winsGemini 2.5 Flash

Benchmark scores

Grouped by capability, sorted by largest gap within each. 12 shared benchmarks.

Coding and Software Engineer

Gemini 3.0 Flash 2/2
BenchmarkGemini 3.0 FlashGemini 2.5 FlashDiff
WeirdML v261.6026 / 52Normal (With Tools)40.9546 / 52Thinking (With Tools, 16K Budget)+20.65
SWE-bench Verified68.7068 / 1145096 / 114+18.70

General Knowledge

Gemini 3.0 Flash 2/2
BenchmarkGemini 3.0 FlashGemini 2.5 FlashDiff
HLE43.5050 / 18111151 / 181+32.50
LiveBench56.3579 / 115Normal (No Tools)47.74101 / 115Thinking High (No Tools)+8.61

Math and Reasoning

Gemini 3.0 Flash 1/2
BenchmarkGemini 3.0 FlashGemini 2.5 FlashDiff
AIME202599.708 / 1077271 / 107+27.70
FrontierMath - Tier 44.2040 / 80Normal (No Tools)4.2040 / 80Normal (No Tools)

Agent Level Benchmark

Gemini 3.0 Flash 1/1
BenchmarkGemini 3.0 FlashGemini 2.5 FlashDiff
BALROG48.103 / 12Normal (With Tools)33.509 / 12Normal (With Tools)+14.60

Claw-style Agent Evaluation

Gemini 3.0 Flash 1/1
BenchmarkGemini 3.0 FlashGemini 2.5 FlashDiff
Pinch Bench85.2017 / 38Thinking (With Tools)70.7032 / 38Thinking (With Tools)+14.50

Common Sense

Gemini 3.0 Flash 1/1
BenchmarkGemini 3.0 FlashGemini 2.5 FlashDiff
SimpleQA68.708 / 4726.9029 / 47+41.80

Commonsense Reasoning

Gemini 3.0 Flash 1/1
BenchmarkGemini 3.0 FlashGemini 2.5 FlashDiff
SimpleBench61.1017 / 67Normal (No Tools)41.2041 / 67Normal (No Tools)+19.90

General Evaluation

Gemini 3.0 Flash 1/1
BenchmarkGemini 3.0 FlashGemini 2.5 FlashDiff
GPQA Diamond90.4036 / 22682.80103 / 226+7.60

Multimodal Understanding

Gemini 3.0 Flash 1/1
BenchmarkGemini 3.0 FlashGemini 2.5 FlashDiff
GeoBench ACW882 / 20Normal (No Tools)769 / 20Normal (No Tools)+12

Specs

FieldGemini 3.0 FlashGemini 2.5 Flash
PublisherGoogle Deep MindGoogle Deep Mind
Release date2025-12-172025-04-17
Model typeChat modelReasoning model
ArchitectureDenseDense
ParametersNot availableNot available
Context length2000K1000K
Max output64K64K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGemini 3.0 FlashGemini 2.5 Flash
Text input$0.5 / 1M tokens$0.3 / 1M tokens
Text output$3 / 1M tokens$2.5 / 1M tokens

Summary

  • Gemini 3.0 Flashleads in:Coding and Software Engineer (2/2), General Knowledge (2/2), Math and Reasoning (1/2), Agent Level Benchmark (1/1), Claw-style Agent Evaluation (1/1), Common Sense (1/1), Commonsense Reasoning (1/1), General Evaluation (1/1), Multimodal Understanding (1/1)

On average across the 12 shared benchmarks, Gemini 3.0 Flash scores 18.21 higher.

Largest single-benchmark gap: SimpleQA — Gemini 3.0 Flash 68.70 vs Gemini 2.5 Flash 26.90 (+41.80).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.