Claude Sonnet 4.6 Benchmark Details
Claude Sonnet 4.6 currently shows benchmark results led by LiveBench (12 / 115, score 75.47), Creative Writing (15 / 99, score 1804.50), SWE-bench Verified (18 / 115, score 79.60). This page also compares it with 3 competitor models and 3 predecessor or same-series models, including performance and pricing views when available. 3 source links are attached for reference.
Benchmark Results
Benchmark Results
General Knowledge
6 evaluationsOther
4 evaluationsCoding and Software Engineer
2 evaluationsWriting and Creative Capabilities
1 evaluationsAI Agent - Tool Usage
3 evaluationsText Embedding
4 evaluationsCompetitor Comparison
Benchmark scores for Claude Sonnet 4.6 compared against top models in its class
12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.
| Benchmark | Claude Sonnet 4.6Current | Claude Opus 4.6 | GPT-5.2 | Gemini 3.0 Pro (Preview 11-2025) |
|---|---|---|---|---|
58.30Thinking Enabled | 66.30Extended Thinking | 54.20Deep Thinking Mode | 45.10Thinking Enabled | |
49.00Thinking Enabled | Tools | 53.00Extended Thinking | Tools | 45.50Deep Thinking Mode | Tools | 45.80Thinking Level · High | Tools | |
75.47Thinking Level · Medium | 76.33Thinking Level · High | 74.84Thinking Level · High | 73.39Thinking Level · High | |
89.90Thinking Enabled | 91.31Extended Thinking | 93.20Deep Thinking Mode | 93.80Thinking Enabled | |
79.60Thinking Enabled | 80.84Extended Thinking | Tools | 80.00Thinking Level · Extra High | Tools | 76.20Thinking Enabled | |
1804.50Standard Mode | 1803.80Standard Mode | 1698.90Standard Mode | -- | |
8.3016K | 22.90Thinking Level · High | 18.80Thinking Level · Extra High | 18.80Standard Mode | |
97.90Thinking Enabled | Tools | 99.25Extended Thinking | Tools | 98.70Thinking Level · Extra High | Tools | 98.00Thinking Level · High | Tools | |
74.70Thinking Enabled | Tools | 84.00Thinking Enabled | Tools | 65.80Thinking Level · Extra High | Tools | 59.20Thinking Level · High | Tools | |
69.50Standard Mode | Tools | 76.80Thinking Level · High | Tools | 67.60Thinking Level · Extra High | Tools | 70.30Standard Mode | Tools | |
72.50Thinking Enabled | Tools | 72.70Extended Thinking | Tools | -- | -- | |
59.10Thinking Enabled | Tools | 65.40Extended Thinking | Tools | -- | 56.90Thinking Level · High | Tools |
Standard API Pricing: Claude Sonnet 4.6 vs. Peer Models
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens
When a context threshold exists, the charted base price only applies within these limits:
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
Claude Sonnet 4.6 | Anthropic | $3 / 1M tokens | $15 / 1M tokens | <= 200K |
Claude Opus 4.6 | Anthropic | $5 / 1M tokens | $25 / 1M tokens | — |
GPT-5.2 | Facebook AI研究实验室 | $1.75 / 1M tokens | $14 / 1M tokens | — |
Gemini 3.0 Pro (Preview 11-2025) | Google Deep Mind | $2 / 1M tokens | $12 / 1M tokens | <= 200000 |
Version History
How each version of the Claude Sonnet 4.6 series stacks up on benchmark tests
12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.
| Benchmark | Claude Sonnet 4.6Current | Claude Sonnet 4.5 | Claude Sonnet 4 | Claude Sonnet 3.7 |
|---|---|---|---|---|
58.30Thinking Enabled | 13.60Thinking Enabled | 5.90Thinking Enabled | -- | |
49.00Thinking Enabled | Tools | 33.60Thinking Enabled | Tools | 9.60Thinking Enabled | 10.30Thinking Enabled | |
75.47Thinking Level · Medium | 68.1964K | 61.2764K | -- | |
89.90Thinking Enabled | 83.40Thinking Enabled | 83.80Deep Thinking Mode | Tools | 77.00Thinking Enabled | |
79.60Thinking Enabled | 82.00Thinking Enabled | Tools | 80.20Thinking Enabled | Tools | 70.30Thinking Enabled | Tools | |
1804.50Standard Mode | 1674.50Standard Mode | 1480.30Standard Mode | -- | |
8.3016K | 4.2032K | 0.00Standard Mode | -- | |
97.90Thinking Enabled | Tools | 98.00Thinking Enabled | Tools | 65.00Thinking Enabled | Tools | 55.00Thinking Enabled | Tools | |
74.70Thinking Enabled | Tools | 24.10Thinking Enabled | Tools | -- | -- | |
69.50Standard Mode | Tools | 59.50Thinking Enabled | Tools | -- | -- | |
72.50Thinking Enabled | Tools | 61.40Thinking Enabled | Tools | 42.20Thinking Enabled | Tools | 28.00Thinking Enabled | Tools | |
59.10Thinking Enabled | Tools | 42.80Thinking Enabled | Tools | -- | -- |
Single-Benchmark Version Trend
Viewing: ARC-AGI-2 · 综合评估
Standard API Pricing Across the Claude Sonnet 4.6 Series
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens
When a context threshold exists, the charted base price only applies within these limits:
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
Claude Sonnet 4.6 | Anthropic | $3 / 1M tokens | $15 / 1M tokens | <= 200K |
Claude Sonnet 4.5 | Anthropic | $3 / 1M tokens | $15 / 1M tokens | <= 200000 |
Claude Sonnet 4 | Anthropic | $3 / 1M tokens | $15 / 1M tokens | <= 200000 |
Claude Sonnet 3.7 | Anthropic | $3 / 1M tokens | $15 / 1M tokens | — |