Claude Sonnet 4.5 Benchmark Details
Claude Sonnet 4.5 currently shows benchmark results led by AIME2025 (1 / 215, score 100), τ²-Bench - Telecom (8 / 264, score 98), MMLU-Pro (8 / 175, score 88). This page also compares it with 2 competitor models and 4 predecessor or same-series models, including performance and pricing views when available. 2 source links are attached for reference.
Benchmark Results
Benchmark Results
Knowledge Exams
6 evaluationsAbstract Generalization
4 evaluationsScientific Reasoning
3 evaluationsRepository Engineering
3 evaluationsAlgorithmic Coding
3 evaluationsMathematics
10 evaluationsAgentic Development
6 evaluationsVisual Understanding
7 evaluationsService Workflows
7 evaluationsInstruction Following
3 evaluationsCross-capability Suites
2 evaluationsTool Orchestration
3 evaluationsML Engineering
2 evaluationsPreference Arenas
2 evaluationsCapability Frontier Metrics
1 evaluationsCode Generation & Editing
1 evaluationsClinical Workflows
4 evaluationsCompetitor Comparison
Benchmark scores for Claude Sonnet 4.5 compared against top models in its class
12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.
| Benchmark | Claude Sonnet 4.5Current | GPT-5.1 | Gemini 2.5-Pro |
|---|---|---|---|
33.60Thinking Enabled | Tools | 42.70Thinking Level · High | Tools | 22.50Thinking Enabled | |
88.00Thinking Enabled | -- | 86.00Standard Mode | |
63.67Thinking Enabled | 72.83Thinking Level · High | 37.00Thinking Enabled | |
13.61Thinking Enabled | 17.64Thinking Level · High | 4.86Thinking Enabled | |
1.10Thinking Enabled | 4.90Thinking Level · High | 2.60Thinking Enabled | |
83.40Thinking Enabled | 88.10Thinking Enabled | 86.40Thinking Enabled | |
43.60Thinking Enabled | 50.80Thinking Level · High | -- | |
82.00Thinking Enabled | Tools | 76.30Thinking Level · High | Tools | 67.20Thinking Enabled | |
1389.00Standard Mode | Tools | -- | 1125.00Standard Mode | Tools | |
71.00Thinking Enabled | 86.80Thinking Level · High | 80.10Thinking Enabled | |
100.00Thinking Enabled | Tools | 94.17Thinking Level · High | Tools | 88.00Thinking Enabled | |
5.20Standard Mode | 26.70Thinking Level · High | Tools | 11.00Standard Mode |
Standard API Pricing: Claude Sonnet 4.5 vs. Peer Models
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens
When a context threshold exists, the charted base price only applies within these limits:
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
Claude Sonnet 4.5 | Anthropic | $3 / 1M tokens | $15 / 1M tokens | <= 200000 |
GPT-5.1 | OpenAI | $1.25 / 1M tokens | $10 / 1M tokens | — |
Gemini 2.5-Pro | Google DeepMind | $1.25 / 1M tokens | $10 / 1M tokens | <= 200000 |
Version History
How each version of the Claude Sonnet 4.5 series stacks up on benchmark tests
12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.
| Benchmark | Claude Sonnet 4.5Current | Claude Sonnet 4 | Claude Sonnet 3.7 | Claude 3.5 Sonnet New | Claude 3.5 Sonnet |
|---|---|---|---|---|---|
33.60Thinking Enabled | Tools | 10.70Thinking Enabled | 10.30Thinking Enabled | -- | 4.08Thinking Level · High | |
88.00Thinking Enabled | 84.00Thinking Enabled | -- | 78.00Standard Mode | 77.64Thinking Level · High | |
63.67Thinking Enabled | 40.00Thinking Enabled | -- | -- | -- | |
13.61Thinking Enabled | 5.93Thinking Enabled | -- | -- | -- | |
1.10Thinking Enabled | 1.10Standard Mode | 0.90Thinking Enabled | -- | -- | |
83.40Thinking Enabled | 83.80Deep Thinking Mode | Tools | 77.00Thinking Enabled | 65.00Standard Mode | 59.40Standard Mode | |
43.60Thinking Enabled | 42.70Thinking Enabled | -- | -- | -- | |
82.00Thinking Enabled | Tools | 80.20Thinking Enabled | Tools | 70.30Thinking Enabled | Tools | 49.00Standard Mode | -- | |
1389.00Standard Mode | Tools | 1223.00Standard Mode | Tools | -- | -- | -- | |
71.00Thinking Enabled | 66.00Thinking Enabled | 47.30Thinking Enabled | 38.70Standard Mode | 38.10Standard Mode | |
100.00Thinking Enabled | Tools | 85.00Deep Thinking Mode | Tools | 56.30Thinking Enabled | -- | -- | |
5.20Standard Mode | 4.10Standard Mode | 4.10Thinking Enabled | 2.10Standard Mode | 1.00Standard Mode |
Single-Benchmark Version Trend
Viewing: HLE · Knowledge Exams
Standard API Pricing Across the Claude Sonnet 4.5 Series
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens
When a context threshold exists, the charted base price only applies within these limits:
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
Claude Sonnet 4.5 | Anthropic | $3 / 1M tokens | $15 / 1M tokens | <= 200000 |
Claude Sonnet 4 | Anthropic | $3 / 1M tokens | $15 / 1M tokens | <= 200000 |
Claude Sonnet 3.7 | Anthropic | $3 / 1M tokens | $15 / 1M tokens | — |
Claude 3.5 Sonnet New | Anthropic | $3 / 1M tokens | $15 / 1M tokens | — |
Claude 3.5 Sonnet | Anthropic | $3 / 1M tokens | $15 / 1M tokens | — |

