DataLearner logo

Gemini 3.8 FlashvsClaude Sonnet 5

Across 3 shared benchmarks, Gemini 3.8 Flash leads overall: Gemini 3.8 Flash wins 3, Claude Sonnet 5 wins 0, with 0 ties and an average score difference of +11.79.

Google Deep Mind
Gemini 3.8 Flash

Google Deep Mind · 2026-09-02 · Multimodal model

Anthropic
Claude Sonnet 5

Anthropic · 2026-06-30 · Multimodal model

Gemini 3.8 Flash3 wins(100%)(0%)0 winsClaude Sonnet 5

Benchmark scores

Grouped by capability, sorted by largest gap within each. 3 shared benchmarks.

AI Agent - Tool Usage

Gemini 3.8 Flash 2/2
BenchmarkGemini 3.8 FlashClaude Sonnet 5Diff
Terminal-Bench 2.189.401 / 48Thinking (With Tools)80.4022 / 48极高强度思考(工具)+9
Terminal-Bench 4.019.109 / 12Thinking (With Tools)12.4211 / 12Max (With Tools)+6.68

Coding and Software Engineer

Gemini 3.8 Flash 1/1
BenchmarkGemini 3.8 FlashClaude Sonnet 5Diff
DeepSWE73.701 / 33Thinking (With Tools)5422 / 33Deep Thinking (With Tools)+19.70

Specs

FieldGemini 3.8 FlashClaude Sonnet 5
PublisherGoogle Deep MindAnthropic
Release date2026-09-022026-06-30
Model typeMultimodal modelMultimodal model
ArchitectureDenseDense
ParametersNot availableNot available
Context length1M1M
Max output64K128K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGemini 3.8 FlashClaude Sonnet 5
Text input$0.75 / 1M tokens$2 / 1M tokens
Text output$3.75 / 1M tokens$10 / 1M tokens
Cache read$0.075 / 1M tokens$0.2 / 1M tokens
Cache writeNot public$2.5 / 1M tokens

Summary

  • Gemini 3.8 Flashleads in:AI Agent - Tool Usage (2/2), Coding and Software Engineer (1/1)

On average across the 3 shared benchmarks, Gemini 3.8 Flash scores 11.79 higher.

Largest single-benchmark gap: DeepSWE — Gemini 3.8 Flash 73.70 vs Claude Sonnet 5 54 (+19.70).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.