DataLearner logo

Gemini 3.0 FlashvsHaiku 4.5

Across 6 shared benchmarks, Gemini 3.0 Flash leads overall: Gemini 3.0 Flash wins 5, Haiku 4.5 wins 1, with 0 ties and an average score difference of +11.62.

Google Deep Mind
Gemini 3.0 Flash

Google Deep Mind · 2025-12-17 · Chat model

Anthropic
Haiku 4.5

Anthropic · 2025-10-15 · Multimodal model

Gemini 3.0 Flash5 wins(83%)(17%)1 winHaiku 4.5

Benchmark scores

Grouped by capability, sorted by largest gap within each. 6 shared benchmarks.

Claw-style Agent Evaluation

Even 2/2
BenchmarkGemini 3.0 FlashHaiku 4.5Diff
Claw Bench85.7015 / 29Thinking (With Tools)89.4011 / 29Thinking (With Tools)-3.70
Pinch Bench85.2017 / 38Thinking (With Tools)8222 / 38Thinking (With Tools)+3.20

AI Agent - Tool Usage

Gemini 3.0 Flash 1/1
BenchmarkGemini 3.0 FlashHaiku 4.5Diff
MCP-Atlas6235 / 43Normal (With Tools)40.2043 / 43Normal (With Tools)+21.80

General Evaluation

Gemini 3.0 Flash 1/1
BenchmarkGemini 3.0 FlashHaiku 4.5Diff
GPQA Diamond81.20215 / 462Normal (No Tools)60.50375 / 462Normal (No Tools)+20.70

General Knowledge

Gemini 3.0 Flash 1/1
BenchmarkGemini 3.0 FlashHaiku 4.5Diff
LiveBench56.3581 / 117Normal (No Tools)45.33105 / 117Normal (No Tools)+11.02

Math and Reasoning

Gemini 3.0 Flash 1/1
BenchmarkGemini 3.0 FlashHaiku 4.5Diff
AIME202555.70146 / 215Normal (No Tools)39168 / 215Normal (No Tools)+16.70

Specs

FieldGemini 3.0 FlashHaiku 4.5
PublisherGoogle Deep MindAnthropic
Release date2025-12-172025-10-15
Model typeChat modelMultimodal model
ArchitectureDenseDense
ParametersNot availableNot available
Context length2000K200K
Max output64K64K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGemini 3.0 FlashHaiku 4.5
Text input$0.5 / 1M tokens$1 / 1M tokens
Text output$3 / 1M tokens$5 / 1M tokens
Cache readNot public$0.1 / 1M tokens
Cache writeNot public$1.25 / 1M tokens

Summary

  • Gemini 3.0 Flashleads in:AI Agent - Tool Usage (1/1), General Evaluation (1/1), General Knowledge (1/1), Math and Reasoning (1/1)
  • Tied in:Claw-style Agent Evaluation

On average across the 6 shared benchmarks, Gemini 3.0 Flash scores 11.62 higher.

Largest single-benchmark gap: MCP-Atlas — Gemini 3.0 Flash 62 vs Haiku 4.5 40.20 (+21.80).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.