DataLearner logo

Gemini 3.6 FlashvsGrok 4.5

Across 8 shared benchmarks, Grok 4.5 leads overall: Gemini 3.6 Flash wins 3, Grok 4.5 wins 5, with 0 ties and an average score difference of -12.07.

Google Deep Mind
Gemini 3.6 Flash

Google Deep Mind · 2026-07-21 · Multimodal model

xAI
Grok 4.5

xAI · 2026-07-08 · Coding model

Gemini 3.6 Flash3 wins(38%)(63%)5 winsGrok 4.5

Benchmark scores

Grouped by capability, sorted by largest gap within each. 8 shared benchmarks.

Coding and Software Engineer

Grok 4.5 2/2
BenchmarkGemini 3.6 FlashGrok 4.5Diff
SWE-Bench Pro - Public58.7016 / 60Thinking (With Tools)64.707 / 60Thinking High (With Tools)-6
DeepSWE4928 / 35Thinking (With Tools)5326 / 35Thinking High (With Tools)-4

Math and Reasoning

Even 2/2
BenchmarkGemini 3.6 FlashGrok 4.5Diff
FrontierMath Tier 4 v221.9525 / 41Thinking High (No Tools)24.3924 / 41Thinking High (No Tools)-2.44
FrontierMath v258.9523 / 58Thinking High (No Tools)57.1924 / 58Thinking High (No Tools)+1.75

AI Agent - Tool Usage

Grok 4.5 1/1
BenchmarkGemini 3.6 FlashGrok 4.5Diff
Terminal-Bench 2.17827 / 49Thinking (With Tools)83.3018 / 49Thinking High (With Tools)-5.30

General Evaluation

Gemini 3.6 Flash 1/1
BenchmarkGemini 3.6 FlashGrok 4.5Diff
GPQA Diamond94.138 / 271Thinking High (No Tools)93.4318 / 271Thinking High (No Tools)+0.69

Productivity Knowledge

Grok 4.5 1/1
BenchmarkGemini 3.6 FlashGrok 4.5Diff
GDPval-AA v21,42118 / 25Thinking (No Tools)1,52615 / 25Thinking High (With Tools)-105

Writing and Creative Capabilities

Gemini 3.6 Flash 1/1
BenchmarkGemini 3.6 FlashGrok 4.5Diff
Creative Writing1,60034 / 99Normal (No Tools)1,57638 / 99Normal (No Tools)+23.70

Specs

FieldGemini 3.6 FlashGrok 4.5
PublisherGoogle Deep MindxAI
Release date2026-07-212026-07-08
Model typeMultimodal modelCoding model
ArchitectureDenseDense
ParametersNot availableNot available
Context length1M500K
Max output64KNot available

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGemini 3.6 FlashGrok 4.5
Text input$1.5 / 1M tokens$2 / 1M tokens
Text output$7.5 / 1M tokens$6 / 1M tokens
Cache read$0.15 / 1M tokens$0.5 / 1M tokens

Summary

  • Gemini 3.6 Flashleads in:General Evaluation (1/1), Writing and Creative Capabilities (1/1)
  • Grok 4.5leads in:Coding and Software Engineer (2/2), AI Agent - Tool Usage (1/1), Productivity Knowledge (1/1)
  • Tied in:Math and Reasoning

On average across the 8 shared benchmarks, Grok 4.5 scores 12.07 higher.

Largest single-benchmark gap: GDPval-AA v2 — Gemini 3.6 Flash 1,421 vs Grok 4.5 1,526 (-105).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.