Claude Sonnet 4.6 Benchmark Details
Claude Sonnet 4.6 currently shows benchmark results led by Terminal Bench Hard (17 / 244, score 53), LiveBench (11 / 117, score 75.32), SWE-bench Verified (18 / 119, score 79.60). This page also compares it with 3 competitor models and 3 predecessor or same-series models, including performance and pricing views when available. 3 source links are attached for reference.
Benchmark Results
Benchmark Results
Abstract Generalization
5 evaluationsScientific Reasoning
6 evaluationsRepository Engineering
2 evaluationsService Workflows
4 evaluationsCross-capability Suites
3 evaluationsAgentic Development
4 evaluationsMemory & Persistence
4 evaluationsTool Orchestration
3 evaluationsCapability Indices
2 evaluationsVisual Understanding
4 evaluationsLegal
4 evaluationsFinance
2 evaluationsCode Generation & Editing
1 evaluationsCompetitor Comparison
Benchmark scores for Claude Sonnet 4.6 compared against top models in its class
12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.
| Benchmark | Claude Sonnet 4.6Current | Claude Opus 4.6 | GPT-5.2 | Gemini 3.0 Pro (Preview 11-2025) |
|---|---|---|---|---|
86.50Thinking Level · High | Tools | 94.00Thinking Level · High | 90.50Deep Thinking Mode | 87.50Thinking Enabled | |
60.42Thinking Level · High | Tools | 69.17Thinking Level · High | 54.20Deep Thinking Mode | 45.10Thinking Enabled | |
49.00Thinking Enabled | Tools | 53.00Extended Thinking | Tools | 45.50Deep Thinking Mode | Tools | 45.80Thinking Level · High | Tools | |
3.10Thinking Level · High | 12.60Thinking Level · High | 11.60Thinking Level · Extra High | 9.10Thinking Level · High | |
89.90Thinking Enabled | 91.31Extended Thinking | 93.20Deep Thinking Mode | 93.80Thinking Enabled | |
79.60Thinking Enabled | 80.84Extended Thinking | Tools | 80.00Thinking Level · Extra High | Tools | 76.20Thinking Enabled | |
1810.40Standard Mode | 1809.10Standard Mode | 1702.90Standard Mode | -- | |
8.3016K | 22.90Thinking Level · High | 18.80Thinking Level · Extra High | 18.80Standard Mode | |
46.58Thinking Level · High | Tools | 51.58Thinking Level · High | Tools | 49.27Thinking Level · Extra High | Tools | -- | |
97.90Thinking Enabled | Tools | 99.25Extended Thinking | Tools | 89.69Thinking Level · High | Tools | 98.00Thinking Enabled | Tools | |
34.40Thinking Level · High | Tools | 27.32Thinking Level · High | Tools | 32.22Thinking Level · High | Tools | -- | |
74.70Thinking Enabled | Tools | 84.00Thinking Enabled | Tools | 65.80Thinking Level · Extra High | Tools | 59.20Thinking Level · High | Tools |
Standard API Pricing: Claude Sonnet 4.6 vs. Peer Models
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens
When a context threshold exists, the charted base price only applies within these limits:
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
Claude Sonnet 4.6 | Anthropic | $3 / 1M tokens | $15 / 1M tokens | — |
Claude Opus 4.6 | Anthropic | $5 / 1M tokens | $25 / 1M tokens | — |
GPT-5.2 | OpenAI | $1.75 / 1M tokens | $14 / 1M tokens | — |
Gemini 3.0 Pro (Preview 11-2025) | Google DeepMind | $2 / 1M tokens | $12 / 1M tokens | <= 200000 |
Version History
How each version of the Claude Sonnet 4.6 series stacks up on benchmark tests
12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.
| Benchmark | Claude Sonnet 4.6Current | Claude Sonnet 4.5 | Claude Sonnet 4 | Claude Sonnet 3.7 |
|---|---|---|---|---|
86.50Thinking Level · High | Tools | 63.67Thinking Enabled | 40.00Thinking Enabled | -- | |
60.42Thinking Level · High | Tools | 13.61Thinking Enabled | 5.93Thinking Enabled | -- | |
49.00Thinking Enabled | Tools | 33.60Thinking Enabled | Tools | 9.60Thinking Enabled | 10.30Thinking Enabled | |
3.10Thinking Level · High | 1.10Thinking Enabled | 1.10Standard Mode | 0.90Thinking Enabled | |
89.90Thinking Enabled | 73.70Standard Mode | 83.80Deep Thinking Mode | Tools | 77.00Thinking Enabled | |
79.60Thinking Enabled | 82.00Thinking Enabled | Tools | 80.20Thinking Enabled | Tools | 70.30Thinking Enabled | Tools | |
1810.40Standard Mode | 1677.60Standard Mode | 1482.80Standard Mode | 1411.70Standard Mode | |
8.3016K | 4.2032K | 0.00Standard Mode | -- | |
46.58Thinking Level · High | Tools | 36.06Thinking Enabled | Tools | 35.00Standard Mode | Tools | -- | |
97.90Thinking Enabled | Tools | 98.00Thinking Enabled | Tools | 65.00Thinking Enabled | Tools | 55.00Thinking Enabled | Tools | |
34.40Thinking Level · High | Tools | 24.50Thinking Enabled | Tools | 16.70Thinking Enabled | Tools | -- | |
74.70Thinking Enabled | Tools | 24.10Thinking Enabled | Tools | -- | -- |
Single-Benchmark Version Trend
Viewing: ARC-AGI-1 · Abstract Generalization
Standard API Pricing Across the Claude Sonnet 4.6 Series
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens
When a context threshold exists, the charted base price only applies within these limits:
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
Claude Sonnet 4.6 | Anthropic | $3 / 1M tokens | $15 / 1M tokens | — |
Claude Sonnet 4.5 | Anthropic | $3 / 1M tokens | $15 / 1M tokens | <= 200000 |
Claude Sonnet 4 | Anthropic | $3 / 1M tokens | $15 / 1M tokens | <= 200000 |
Claude Sonnet 3.7 | Anthropic | $3 / 1M tokens | $15 / 1M tokens | — |


