DataLearner logo

Claude Sonnet 4.5vsGemini 2.5-Pro

Across 24 shared benchmarks, Claude Sonnet 4.5 leads overall: Claude Sonnet 4.5 wins 14, Gemini 2.5-Pro wins 8, with 2 ties and an average score difference of +16.50.

Anthropic
Claude Sonnet 4.5

Anthropic · 2025-09-30 · Chat model

Google Deep Mind
Gemini 2.5-Pro

Google Deep Mind · 2025-06-05 · Reasoning model

Claude Sonnet 4.514 wins(58%)Ties2(33%)8 winsGemini 2.5-Pro

Benchmark scores

Grouped by capability, sorted by largest gap within each. 24 shared benchmarks.

General Knowledge

Claude Sonnet 4.5 4/6
BenchmarkClaude Sonnet 4.5Gemini 2.5-ProDiff
ARC-AGI63.7035 / 683750 / 68+26.70
HLE33.6080 / 17221.60112 / 172+12
ARC-AGI-213.6038 / 624.9047 / 62+8.70
LiveBench53.6983 / 115Normal (No Tools)58.3376 / 115Thinking High (No Tools)-4.64
GPQA Diamond83.4063 / 18786.4045 / 187-3
MMLU Pro887 / 1328621 / 132+2

Math and Reasoning

Gemini 2.5-Pro 3/5
BenchmarkClaude Sonnet 4.5Gemini 2.5-ProDiff
IMO-ProofBench27.108 / 1655.203 / 16-28.10
IMO-ProofBench Advanced4.806 / 817.604 / 8-12.80
AIME20251001 / 1078844 / 107+12
FrontierMath5.2038 / 601123 / 60-5.80
FrontierMath - Tier 42.1056 / 80Normal (No Tools)2.1056 / 80Normal (No Tools)

Coding and Software Engineer

Claude Sonnet 4.5 2/3
BenchmarkClaude Sonnet 4.5Gemini 2.5-ProDiff
CodeClash1,3891 / 8Normal (With Tools)1,1256 / 8Normal (With Tools)+264
SWE-bench Verified828 / 11267.2072 / 112+14.80
LiveCodeBench7148 / 12377.1034 / 123-6.10

Agent Level Benchmark

Claude Sonnet 4.5 2/2
BenchmarkClaude Sonnet 4.5Gemini 2.5-ProDiff
τ²-Bench - Telecom985 / 355432 / 35+44
Terminal Bench Hard338 / 132512 / 13+8

AI Agent - Tool Usage

Claude Sonnet 4.5 2/2
BenchmarkClaude Sonnet 4.5Gemini 2.5-ProDiff
Terminal-Bench503 / 3525.3028 / 35+24.70
Terminal Bench 2.042.8042 / 4732.6047 / 47+10.20

AI Agent - Information Search

Claude Sonnet 4.5 1/1
BenchmarkClaude Sonnet 4.5Gemini 2.5-ProDiff
BrowseComp24.1051 / 537.8052 / 53+16.30

Commonsense Reasoning

Gemini 2.5-Pro 1/1
BenchmarkClaude Sonnet 4.5Gemini 2.5-ProDiff
Simple Bench54.3022 / 63Normal (No Tools)62.4011 / 63Thinking (No Tools)-8.10

Instruction Following

Claude Sonnet 4.5 1/1
BenchmarkClaude Sonnet 4.5Gemini 2.5-ProDiff
IF Bench57.3022 / 304929 / 30+8.30

Long Context

Even 1/1
BenchmarkClaude Sonnet 4.5Gemini 2.5-ProDiff
AA-LCR6610 / 156610 / 15

Multimodal Understanding

Gemini 2.5-Pro 1/1
BenchmarkClaude Sonnet 4.5Gemini 2.5-ProDiff
MMMU77.8015 / 298210 / 29-4.20

Productivity Knowledge

Claude Sonnet 4.5 1/1
BenchmarkClaude Sonnet 4.5Gemini 2.5-ProDiff
GDPval-AA3916 / 212221 / 21+17

Specs

FieldClaude Sonnet 4.5Gemini 2.5-Pro
PublisherAnthropicGoogle Deep Mind
Release date2025-09-302025-06-05
Model typeChat modelReasoning model
ArchitectureDenseDense
ParametersNot availableNot available
Context length1000K1000K
Max output64K64K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemClaude Sonnet 4.5Gemini 2.5-Pro
Text input$3 / 1M tokens$1.25 / 1M tokens
Text output$15 / 1M tokens$10 / 1M tokens
Cache read$0.3 / 1M tokensNot public
Cache write$3.75 / 1M tokensNot public

Summary

  • Claude Sonnet 4.5leads in:General Knowledge (4/6), Coding and Software Engineer (2/3), Agent Level Benchmark (2/2), AI Agent - Tool Usage (2/2), AI Agent - Information Search (1/1), Instruction Following (1/1), Productivity Knowledge (1/1)
  • Gemini 2.5-Proleads in:Math and Reasoning (3/5), Commonsense Reasoning (1/1), Multimodal Understanding (1/1)
  • Tied in:Long Context

On average across the 24 shared benchmarks, Claude Sonnet 4.5 scores 16.50 higher.

Largest single-benchmark gap: CodeClash — Claude Sonnet 4.5 1,389 vs Gemini 2.5-Pro 1,125 (+264).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.