DataLearner logo

Gemini 3.7 FlashvsGemini 3.0 Flash

Across 3 shared benchmarks, Gemini 3.7 Flash leads overall: Gemini 3.7 Flash wins 3, Gemini 3.0 Flash wins 0, with 0 ties and an average score difference of +27.74.

Google Deep Mind
Gemini 3.7 Flash

Google Deep Mind · 2026-08-13 · Multimodal model

Google Deep Mind
Gemini 3.0 Flash

Google Deep Mind · 2025-12-17 · Chat model

Gemini 3.7 Flash3 wins(100%)(0%)0 winsGemini 3.0 Flash

Benchmark scores

Grouped by capability, sorted by largest gap within each. 3 shared benchmarks.

AI Agent - Tool Usage

Gemini 3.7 Flash 1/1
BenchmarkGemini 3.7 FlashGemini 3.0 FlashDiff
Terminal-Bench 2.185.809 / 47Thinking (With Tools)5843 / 47Thinking High (With Tools)+27.80

General Evaluation

Gemini 3.7 Flash 1/1
BenchmarkGemini 3.7 FlashGemini 3.0 FlashDiff
GPQA Diamond94.821 / 270Thinking High (No Tools)90.4042 / 270Thinking (No Tools)+4.42

General Knowledge

Gemini 3.7 Flash 1/1
BenchmarkGemini 3.7 FlashGemini 3.0 FlashDiff
ARC-AGI-284.6014 / 84Thinking High (No Tools)33.6051 / 84Thinking (No Tools)+51

Specs

FieldGemini 3.7 FlashGemini 3.0 Flash
PublisherGoogle Deep MindGoogle Deep Mind
Release date2026-08-132025-12-17
Model typeMultimodal modelChat model
ArchitectureDenseDense
ParametersNot availableNot available
Context length1M2000K
Max output64K64K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGemini 3.7 FlashGemini 3.0 Flash
Text input$0.75 / 1M tokens$0.5 / 1M tokens
Text output$3.75 / 1M tokens$3 / 1M tokens
Cache read$0.075 / 1M tokensNot public

Summary

  • Gemini 3.7 Flashleads in:AI Agent - Tool Usage (1/1), General Evaluation (1/1), General Knowledge (1/1)

On average across the 3 shared benchmarks, Gemini 3.7 Flash scores 27.74 higher.

Largest single-benchmark gap: ARC-AGI-2 — Gemini 3.7 Flash 84.60 vs Gemini 3.0 Flash 33.60 (+51).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.