DataLearner logo

Claude Sonnet 5vsGemini 3.5 Flash

Across 8 shared benchmarks, Claude Sonnet 5 leads overall: Claude Sonnet 5 wins 6, Gemini 3.5 Flash wins 2, with 0 ties and an average score difference of +3.17.

Anthropic
Claude Sonnet 5

Anthropic · 2026-06-30 · Multimodal model

Google Deep Mind
Gemini 3.5 Flash

Google Deep Mind · 2026-06-20 · Multimodal model

Claude Sonnet 56 wins(75%)(25%)2 winsGemini 3.5 Flash

Benchmark scores

Grouped by capability, sorted by largest gap within each. 8 shared benchmarks.

AI Agent - Tool Usage

Claude Sonnet 5 2/2
BenchmarkClaude Sonnet 5Gemini 3.5 FlashDiff
Terminal-Bench 2.180.4018 / 44极高强度思考(工具)76.2023 / 44Thinking High (With Tools)+4.20
OSWorld-Verified81.206 / 26极高强度思考(工具)78.4010 / 26Thinking High (With Tools)+2.80

Math and Reasoning

Claude Sonnet 5 2/2
BenchmarkClaude Sonnet 5Gemini 3.5 FlashDiff
FrontierMath v265.6117 / 34最高(无工具)62.8118 / 34Thinking High (No Tools)+2.81
FrontierMath Tier 4 v229.2716 / 34最高(无工具)26.8318 / 34Thinking High (No Tools)+2.44

Coding and Software Engineer

Claude Sonnet 5 1/1
BenchmarkClaude Sonnet 5Gemini 3.5 FlashDiff
DeepSWE5416 / 27Deep Thinking (With Tools)3723 / 27Thinking Medium (With Tools)+17

Commonsense Reasoning

Gemini 3.5 Flash 1/1
BenchmarkClaude Sonnet 5Gemini 3.5 FlashDiff
SimpleBench57.9023 / 67Normal (No Tools)76.704 / 67Normal (No Tools)-18.80

General Evaluation

Gemini 3.5 Flash 1/1
BenchmarkClaude Sonnet 5Gemini 3.5 FlashDiff
GPQA Diamond90.5333 / 226极高强度思考(无工具)92.8019 / 226Thinking High (No Tools)-2.27

General Knowledge

Claude Sonnet 5 1/1
BenchmarkClaude Sonnet 5Gemini 3.5 FlashDiff
HLE57.409 / 181极高强度思考(工具)40.2066 / 181Thinking High (With Tools)+17.20

Specs

FieldClaude Sonnet 5Gemini 3.5 Flash
PublisherAnthropicGoogle Deep Mind
Release date2026-06-302026-06-20
Model typeMultimodal modelMultimodal model
ArchitectureDenseDense
ParametersNot availableNot available
Context length1M1M
Max output128K64K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemClaude Sonnet 5Gemini 3.5 Flash
Text input$2 / 1M tokens$1.5 / 1M tokens
Text output$10 / 1M tokens$9 / 1M tokens
Cache read$0.3 / 1M tokens$0.15 / 1M tokens
Cache write$3.75 / 1M tokensNot public

Summary

  • Claude Sonnet 5leads in:AI Agent - Tool Usage (2/2), Math and Reasoning (2/2), Coding and Software Engineer (1/1), General Knowledge (1/1)
  • Gemini 3.5 Flashleads in:Commonsense Reasoning (1/1), General Evaluation (1/1)

On average across the 8 shared benchmarks, Claude Sonnet 5 scores 3.17 higher.

Largest single-benchmark gap: SimpleBench — Claude Sonnet 5 57.90 vs Gemini 3.5 Flash 76.70 (-18.80).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.