Claude Sonnet 5.5 Benchmark Details
Claude Sonnet 5.5 currently shows benchmark results led by Terminal-Bench 4.0 (1 / 97, score 70.60), HLE (4 / 235, score 64.50), CursorBench 4.0 (2 / 47, score 55.50). This page also compares it with 3 competitor models and 3 predecessor or same-series models, including performance and pricing views when available.
Benchmark Results
Benchmark Results
Agentic Development
2 evaluationsScientific Computing
1 evaluationsCode Generation & Editing
2 evaluationsCompetitor Comparison
Benchmark scores for Claude Sonnet 5.5 compared against top models in its class
6 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.
| Benchmark | Claude Sonnet 5.5Current | GPT-6 Sol | Muse Spark 1.3 | DeepSeek-V4.1-Flash |
|---|---|---|---|---|
64.50Thinking Level · High | Tools | -- | -- | 63.90Thinking Level · High | Tools | |
44.70Thinking Level · High | Tools | 33.20Thinking Level · Extra High | Tools | 49.40Thinking Level · High | Tools | 54.80Thinking Level · High | Tools | |
61.60Thinking Level · High | -- | -- | 78.90Thinking Level · High | Tools | |
55.50Thinking Level · High | Tools | -- | 41.60Thinking Level · High | Tools | -- | |
70.60Thinking Level · High | Tools | 43.94Thinking Level · High | Tools | 33.30Thinking Level · High | Tools | 26.80Thinking Level · High | Tools | |
52.10Thinking Level · Extra High | Tools | 49.30Thinking Level · High | Tools | -- | -- |
Standard API Pricing: Claude Sonnet 5.5 vs. Peer Models
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens
When a context threshold exists, the charted base price only applies within these limits:
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
Claude Sonnet 5.5 | Anthropic | $2 / 1M tokens | $10 / 1M tokens | — |
GPT-6 Sol | OpenAI | $2 / 1M tokens | $10 / 1M tokens | <= 272000 |
Muse Spark 1.3 | Facebook AI研究实验室 | $1.25 / 1M tokens | $4.25 / 1M tokens | — |
Version History
How each version of the Claude Sonnet 5.5 series stacks up on benchmark tests
3 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.
| Benchmark | Claude Sonnet 5.5Current | Claude Sonnet 5 | Claude Sonnet 4.6 | Claude Sonnet 4.5 |
|---|---|---|---|---|
64.50Thinking Level · High | Tools | 57.40Thinking Level · Extra High | Tools | 49.00Thinking Enabled | Tools | 33.60Thinking Enabled | Tools | |
55.50Thinking Level · High | Tools | 34.10Thinking Level · High | Tools | -- | -- | |
70.60Thinking Level · High | Tools | 12.42Thinking Level · High | Tools | 3.00Thinking Level · High | Tools | -- |
Single-Benchmark Version Trend
Viewing: HLE · Knowledge Exams
Standard API Pricing Across the Claude Sonnet 5.5 Series
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens
When a context threshold exists, the charted base price only applies within these limits:
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
Claude Sonnet 5.5 | Anthropic | $2 / 1M tokens | $10 / 1M tokens | — |
Claude Sonnet 5 | Anthropic | $2 / 1M tokens | $10 / 1M tokens | — |
Claude Sonnet 4.6 | Anthropic | $3 / 1M tokens | $15 / 1M tokens | — |
Claude Sonnet 4.5 | Anthropic | $3 / 1M tokens | $15 / 1M tokens | <= 200000 |


