DataLearner logo

Claude Sonnet 5.5 Benchmark Details

Claude Sonnet 5.5 currently shows benchmark results led by Terminal-Bench 4.0 (1 / 97, score 70.60), HLE (4 / 235, score 64.50), CursorBench 4.0 (2 / 47, score 55.50). This page also compares it with 3 competitor models and 3 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

Claude Sonnet 5.5

Benchmark Results

Thinking
Tool usage

Knowledge Exams

2 evaluations
Benchmark / mode
Score
Rank/total
HLE
Thinking Level · Max
56.90
19 / 235
HLE
Thinking Level · MaxTools
64.50
4 / 235

Office & Business

1 evaluations
Benchmark / mode
Score
Rank/total
AutomationBench
Thinking Level · MaxTools
44.70
11 / 24

Documents & Charts

1 evaluations
Benchmark / mode
Score
Rank/total
Chartography
Thinking Level · Max
61.60
9 / 10

Agentic Development

2 evaluations
Benchmark / mode
Score
Rank/total
Terminal-Bench 4.0
Thinking Level · MaxTools
70.60
1 / 97
CursorBench 4.0
Thinking Level · MaxTools
55.50
2 / 47

Scientific Computing

1 evaluations
Benchmark / mode
Score
Rank/total
Terminal-Bench-Science 0.1
Thinking Level · MaxTools
59.90
2 / 13

Code Generation & Editing

2 evaluations
Benchmark / mode
Score
Rank/total
FrontierCode 1.1 Main
Thinking Level · MaxTools
46.20
10 / 11
FrontierCode 1.1 Main
Thinking Level · Extra HighTools
52.10
5 / 11

Competitor Comparison

Benchmark scores for Claude Sonnet 5.5 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

6 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkClaude Sonnet 5.5CurrentGPT-6 SolMuse Spark 1.3DeepSeek-V4.1-Flash
HLE
Accuracy
Knowledge Exams
64.50Thinking Level · High | Tools
--
--
63.90Thinking Level · High | Tools
AutomationBench
Pass Rate
Office & Business
44.70Thinking Level · High | Tools
33.20Thinking Level · Extra High | Tools
49.40Thinking Level · High | Tools
54.80Thinking Level · High | Tools
Chartography
Score
Documents & Charts
61.60Thinking Level · High
--
--
78.90Thinking Level · High | Tools
CursorBench 4.0
Task score (%)
Agentic Development
55.50Thinking Level · High | Tools
--
41.60Thinking Level · High | Tools
--
Terminal-Bench 4.0
Resolution rate (%)
Agentic Development
70.60Thinking Level · High | Tools
43.94Thinking Level · High | Tools
33.30Thinking Level · High | Tools
26.80Thinking Level · High | Tools
FrontierCode 1.1 Main
Score (%)
Code Generation & Editing
52.10Thinking Level · Extra High | Tools
49.30Thinking Level · High | Tools
--
--

Standard API Pricing: Claude Sonnet 5.5 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

GPT-6 Sol: Base price applies to <= 272000
ModelSupplierStandard inputStandard outputBase price applies to
Claude Sonnet 5.5
Anthropic$2 / 1M tokens$10 / 1M tokens—
GPT-6 Sol
OpenAI$2 / 1M tokens$10 / 1M tokens<= 272000
Muse Spark 1.3
Facebook AI研究实验室$1.25 / 1M tokens$4.25 / 1M tokens—

Version History

How each version of the Claude Sonnet 5.5 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

3 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkClaude Sonnet 5.5CurrentClaude Sonnet 5Claude Sonnet 4.6Claude Sonnet 4.5
HLE
Accuracy
Knowledge Exams
64.50Thinking Level · High | Tools
57.40Thinking Level · Extra High | Tools
49.00Thinking Enabled | Tools
33.60Thinking Enabled | Tools
CursorBench 4.0
Task score (%)
Agentic Development
55.50Thinking Level · High | Tools
34.10Thinking Level · High | Tools
--
--
Terminal-Bench 4.0
Resolution rate (%)
Agentic Development
70.60Thinking Level · High | Tools
12.42Thinking Level · High | Tools
3.00Thinking Level · High | Tools
--

Single-Benchmark Version Trend

Viewing: HLE · Knowledge Exams

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Claude Sonnet 5.5 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Claude Sonnet 4.5: Base price applies to <= 200000
ModelSupplierStandard inputStandard outputBase price applies to
Claude Sonnet 5.5
Anthropic$2 / 1M tokens$10 / 1M tokens—
Claude Sonnet 5
Anthropic$2 / 1M tokens$10 / 1M tokens—
Claude Sonnet 4.6
Anthropic$3 / 1M tokens$15 / 1M tokens—
Claude Sonnet 4.5
Anthropic$3 / 1M tokens$15 / 1M tokens<= 200000
Claude Sonnet 5.5 Benchmark Results Analysis & Model Comparisons | DataLearnerAI