DataLearner logo

Gemini 3.6 FlashvsClaude Sonnet 5

Across 3 shared benchmarks, Claude Sonnet 5 leads overall: Gemini 3.6 Flash wins 1, Claude Sonnet 5 wins 2, with 0 ties and an average score difference of -1.87.

Google Deep Mind
Gemini 3.6 Flash

Google Deep Mind · 2026-07-21 · Multimodal model

Anthropic
Claude Sonnet 5

Anthropic · 2026-06-30 · Multimodal model

Gemini 3.6 Flash1 win(33%)(67%)2 winsClaude Sonnet 5

Benchmark scores

Grouped by capability, sorted by largest gap within each. 3 shared benchmarks.

AI Agent - Tool Usage

Even 2/2
BenchmarkGemini 3.6 FlashClaude Sonnet 5Diff
TerminalBench 2.17813 / 27Thinking (With Tools)80.4010 / 27极高强度思考(工具)-2.40
OSWorld-Verified833 / 23Thinking (With Tools)81.204 / 23极高强度思考(工具)+1.80

Coding and Software Engineer

Claude Sonnet 5 1/1
BenchmarkGemini 3.6 FlashClaude Sonnet 5Diff
DeepSWE4912 / 18Thinking (With Tools)548 / 18Deep Thinking (With Tools)-5

Specs

FieldGemini 3.6 FlashClaude Sonnet 5
PublisherGoogle Deep MindAnthropic
Release date2026-07-212026-06-30
Model typeMultimodal modelMultimodal model
ArchitectureDenseDense
ParametersNot availableNot available
Context length1M1M
Max output64K128K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGemini 3.6 FlashClaude Sonnet 5
Text input$1.5 / 1M tokens$2 / 1M tokens
Text output$7.5 / 1M tokens$10 / 1M tokens
Cache read$0.15 / 1M tokens$0.3 / 1M tokens
Cache writeNot public$3.75 / 1M tokens

Summary

  • Claude Sonnet 5leads in:Coding and Software Engineer (1/1)
  • Tied in:AI Agent - Tool Usage

On average across the 3 shared benchmarks, Claude Sonnet 5 scores 1.87 higher.

Largest single-benchmark gap: DeepSWE — Gemini 3.6 Flash 49 vs Claude Sonnet 5 54 (-5).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.