DataLearner logo

Claude Sonnet 4.5vsGemini 2.5-Pro

Across 17 shared benchmarks, Gemini 2.5-Pro leads overall: Claude Sonnet 4.5 wins 7, Gemini 2.5-Pro wins 9, with 1 ties and an average score difference of +30.39.

Anthropic
Claude Sonnet 4.5

Anthropic · 2025-09-30 · Chat model

Google Deep Mind
Gemini 2.5-Pro

Google Deep Mind · 2025-06-05 · Reasoning model

Claude Sonnet 4.57 wins(41%)Ties1(53%)9 winsGemini 2.5-Pro

Benchmark scores

Grouped by capability, sorted by largest gap within each. 17 shared benchmarks.

Coding and Software Engineer

Even 4/4
BenchmarkClaude Sonnet 4.5Gemini 2.5-ProDiff
CodeClash1,3891 / 8Normal (With Tools)1,1256 / 8Normal (With Tools)+264
LiveCodeBench59136 / 250Normal (No Tools)77.1061 / 250Normal (No Tools)-18.10
GSO14.7012 / 21Normal (With Tools)3.9218 / 21Normal (With Tools)+10.78
SciCode45.7086 / 130Thinking (No Tools)46.3083 / 130Thinking (No Tools)-0.60

Math and Reasoning

Gemini 2.5-Pro 3/4
BenchmarkClaude Sonnet 4.5Gemini 2.5-ProDiff
IMO-ProofBench27.108 / 16Thinking (No Tools)55.203 / 16Thinking (No Tools)-28.10
IMO-ProofBench Advanced4.8019 / 24Thinking (No Tools)17.6014 / 24Thinking (No Tools)-12.80
FrontierMath5.2038 / 60Normal (No Tools)1123 / 60Normal (No Tools)-5.80
FrontierMath - Tier 42.1056 / 80Normal (No Tools)2.1056 / 80Normal (No Tools)

Multimodal Understanding

Gemini 2.5-Pro 3/3
BenchmarkClaude Sonnet 4.5Gemini 2.5-ProDiff
VPCT3820 / 24Normal (No Tools)46.4012 / 24Normal (No Tools)-8.40
GDP.pdf5.2095 / 118Thinking (No Tools)10.2081 / 118Thinking (No Tools)-5
MMMU77.8023 / 74Thinking (No Tools)8213 / 74Thinking (No Tools)-4.20

AI Agent - Tool Usage

Claude Sonnet 4.5 2/2
BenchmarkClaude Sonnet 4.5Gemini 2.5-ProDiff
Terminal-Bench 2.155.80120 / 192Thinking (With Tools)28.50158 / 192Thinking (With Tools)+27.30
Terminal Bench 2.042.8043 / 48Thinking (With Tools)32.6048 / 48Thinking (With Tools)+10.20

AI Agent - Information Search

Claude Sonnet 4.5 1/1
BenchmarkClaude Sonnet 4.5Gemini 2.5-ProDiff
BrowseComp24.1055 / 57Thinking (With Tools)7.8056 / 57Thinking (With Tools)+16.30

General Knowledge

Gemini 2.5-Pro 1/1
BenchmarkClaude Sonnet 4.5Gemini 2.5-ProDiff
CritPt1.10147 / 200Thinking (No Tools)2.60121 / 200Thinking (No Tools)-1.50

Productivity Knowledge

Claude Sonnet 4.5 1/1
BenchmarkClaude Sonnet 4.5Gemini 2.5-ProDiff
GDPval-AA3910 / 15Thinking (No Tools)2215 / 15Thinking (No Tools)+17

Writing and Creative Capabilities

Claude Sonnet 4.5 1/1
BenchmarkClaude Sonnet 4.5Gemini 2.5-ProDiff
Creative Writing1,67430 / 106Normal (No Tools)1,41958 / 106Normal (No Tools)+255.50

Specs

FieldClaude Sonnet 4.5Gemini 2.5-Pro
PublisherAnthropicGoogle Deep Mind
Release date2025-09-302025-06-05
Model typeChat modelReasoning model
ArchitectureDenseDense
ParametersNot availableNot available
Context length1000K1000K
Max output64K64K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemClaude Sonnet 4.5Gemini 2.5-Pro
Text input$3 / 1M tokens$1.25 / 1M tokens
Text output$15 / 1M tokens$10 / 1M tokens
Cache read$0.3 / 1M tokens$0.125 / 1M tokens
Cache write$3.75 / 1M tokensNot public

Summary

  • Claude Sonnet 4.5leads in:AI Agent - Tool Usage (2/2), AI Agent - Information Search (1/1), Productivity Knowledge (1/1), Writing and Creative Capabilities (1/1)
  • Gemini 2.5-Proleads in:Math and Reasoning (3/4), Multimodal Understanding (3/3), General Knowledge (1/1)
  • Tied in:Coding and Software Engineer

On average across the 17 shared benchmarks, Claude Sonnet 4.5 scores 30.39 higher.

Largest single-benchmark gap: CodeClash — Claude Sonnet 4.5 1,389 vs Gemini 2.5-Pro 1,125 (+264).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.