Claude Sonnet 4.5vsGemini 2.5-Pro
Across 17 shared benchmarks, Gemini 2.5-Pro leads overall: Claude Sonnet 4.5 wins 7, Gemini 2.5-Pro wins 9, with 1 ties and an average score difference of +30.39.
Claude Sonnet 4.5
Anthropic · 2025-09-30 · Chat model
Gemini 2.5-Pro
Google Deep Mind · 2025-06-05 · Reasoning model
Claude Sonnet 4.57 wins(41%)Ties1(53%)9 winsGemini 2.5-Pro
Benchmark scores
Grouped by capability, sorted by largest gap within each. 17 shared benchmarks.
Coding and Software Engineer
Even 4/4| Benchmark | Claude Sonnet 4.5 | Gemini 2.5-Pro | Diff |
|---|---|---|---|
| CodeClash | 1,3891 / 8Normal (With Tools) | 1,1256 / 8Normal (With Tools) | +264 |
| LiveCodeBench | 59136 / 250Normal (No Tools) | 77.1061 / 250Normal (No Tools) | -18.10 |
| GSO | 14.7012 / 21Normal (With Tools) | 3.9218 / 21Normal (With Tools) | +10.78 |
| SciCode | 45.7086 / 130Thinking (No Tools) | 46.3083 / 130Thinking (No Tools) | -0.60 |
Math and Reasoning
Gemini 2.5-Pro 3/4| Benchmark | Claude Sonnet 4.5 | Gemini 2.5-Pro | Diff |
|---|---|---|---|
| IMO-ProofBench | 27.108 / 16Thinking (No Tools) | 55.203 / 16Thinking (No Tools) | -28.10 |
| IMO-ProofBench Advanced | 4.8019 / 24Thinking (No Tools) | 17.6014 / 24Thinking (No Tools) | -12.80 |
| FrontierMath | 5.2038 / 60Normal (No Tools) | 1123 / 60Normal (No Tools) | -5.80 |
| FrontierMath - Tier 4 | 2.1056 / 80Normal (No Tools) | 2.1056 / 80Normal (No Tools) | — |
Multimodal Understanding
Gemini 2.5-Pro 3/3| Benchmark | Claude Sonnet 4.5 | Gemini 2.5-Pro | Diff |
|---|---|---|---|
| VPCT | 3820 / 24Normal (No Tools) | 46.4012 / 24Normal (No Tools) | -8.40 |
| GDP.pdf | 5.2095 / 118Thinking (No Tools) | 10.2081 / 118Thinking (No Tools) | -5 |
| MMMU | 77.8023 / 74Thinking (No Tools) | 8213 / 74Thinking (No Tools) | -4.20 |
AI Agent - Tool Usage
Claude Sonnet 4.5 2/2| Benchmark | Claude Sonnet 4.5 | Gemini 2.5-Pro | Diff |
|---|---|---|---|
| Terminal-Bench 2.1 | 55.80120 / 192Thinking (With Tools) | 28.50158 / 192Thinking (With Tools) | +27.30 |
| Terminal Bench 2.0 | 42.8043 / 48Thinking (With Tools) | 32.6048 / 48Thinking (With Tools) | +10.20 |
AI Agent - Information Search
Claude Sonnet 4.5 1/1| Benchmark | Claude Sonnet 4.5 | Gemini 2.5-Pro | Diff |
|---|---|---|---|
| BrowseComp | 24.1055 / 57Thinking (With Tools) | 7.8056 / 57Thinking (With Tools) | +16.30 |
General Knowledge
Gemini 2.5-Pro 1/1| Benchmark | Claude Sonnet 4.5 | Gemini 2.5-Pro | Diff |
|---|---|---|---|
| CritPt | 1.10147 / 200Thinking (No Tools) | 2.60121 / 200Thinking (No Tools) | -1.50 |
Productivity Knowledge
Claude Sonnet 4.5 1/1| Benchmark | Claude Sonnet 4.5 | Gemini 2.5-Pro | Diff |
|---|---|---|---|
| GDPval-AA | 3910 / 15Thinking (No Tools) | 2215 / 15Thinking (No Tools) | +17 |
Writing and Creative Capabilities
Claude Sonnet 4.5 1/1| Benchmark | Claude Sonnet 4.5 | Gemini 2.5-Pro | Diff |
|---|---|---|---|
| Creative Writing | 1,67430 / 106Normal (No Tools) | 1,41958 / 106Normal (No Tools) | +255.50 |
Specs
| Field | Claude Sonnet 4.5 | Gemini 2.5-Pro |
|---|---|---|
| Publisher | Anthropic | Google Deep Mind |
| Release date | 2025-09-30 | 2025-06-05 |
| Model type | Chat model | Reasoning model |
| Architecture | Dense | Dense |
| Parameters | Not available | Not available |
| Context length | 1000K | 1000K |
| Max output | 64K | 64K |
API pricing
Prices use DataLearner records when available; missing fields are not inferred.
| Item | Claude Sonnet 4.5 | Gemini 2.5-Pro |
|---|---|---|
| Text input | $3 / 1M tokens | $1.25 / 1M tokens |
| Text output | $15 / 1M tokens | $10 / 1M tokens |
| Cache read | $0.3 / 1M tokens | $0.125 / 1M tokens |
| Cache write | $3.75 / 1M tokens | Not public |
Summary
- Claude Sonnet 4.5leads in:AI Agent - Tool Usage (2/2), AI Agent - Information Search (1/1), Productivity Knowledge (1/1), Writing and Creative Capabilities (1/1)
- Gemini 2.5-Proleads in:Math and Reasoning (3/4), Multimodal Understanding (3/3), General Knowledge (1/1)
- Tied in:Coding and Software Engineer
On average across the 17 shared benchmarks, Claude Sonnet 4.5 scores 30.39 higher.
Largest single-benchmark gap: CodeClash — Claude Sonnet 4.5 1,389 vs Gemini 2.5-Pro 1,125 (+264).
Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.