DataLearner logo

Claude Sonnet 4.6 Benchmark Details

Claude Sonnet 4.6 currently shows benchmark results led by LiveBench (12 / 115, score 75.47), Creative Writing (15 / 99, score 1804.50), SWE-bench Verified (18 / 115, score 79.60). This page also compares it with 3 competitor models and 3 predecessor or same-series models, including performance and pricing views when available. 3 source links are attached for reference.

Benchmark Results

Claude Sonnet 4.6

Benchmark Results

Thinking
Tool usage

General Knowledge

6 evaluations
Benchmark / mode
Score
Rank/total
LiveBench
Thinking Level · Low
70.44
36 / 115
LiveBench
Thinking Level · Medium
75.47
12 / 115
LiveBench
Thinking Level · High
75.32
15 / 115
ARC-AGI-2
Thinking Mode
58.30
41 / 85
HLE
Thinking Mode
33.20
93 / 189
HLE
Thinking ModeTools
49
39 / 189

Other

4 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Level · Medium
83.33
114 / 270
GPQA Diamond
Thinking Mode
89.90
49 / 270
GPQA Diamond
Thinking Level · High
83.33
114 / 270
GPQA Diamond
Thinking Level · Max
78.79
152 / 270

Coding and Software Engineer

2 evaluations
Benchmark / mode
Score
Rank/total
SWE-bench Verified
Thinking Mode
79.60
18 / 115
DeepSWE
Thinking Level · HighTools
30
32 / 34

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1804.50
15 / 99

Math and Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
8.30
34 / 80

Agent Level Benchmark

1 evaluations
Benchmark / mode
Score
Rank/total
τ²-Bench - Telecom
Thinking ModeTools
97.90
9 / 35

AI Agent - Information Search

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
Thinking ModeTools
74.70
30 / 58

AI Agent - Tool Usage

3 evaluations
Benchmark / mode
Score
Rank/total
OSWorld-Verified
Thinking ModeTools
72.50
17 / 26
MCP-Atlas
Standard ModeTools
69.50
29 / 41
Terminal Bench 2.0
Thinking ModeTools
59.10
22 / 48

Text Embedding

4 evaluations
Benchmark / mode
Score
Rank/total
Context Arena
Standard Mode
55.18
80 / 126
Context Arena
Thinking Level · Low
82.80
30 / 126
Context Arena
Thinking Level · Medium
82.56
31 / 126
Context Arena
Thinking Level · High
82.82
29 / 126

Productivity Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA
Thinking Mode
57
11 / 21

Long Context

1 evaluations
Benchmark / mode
Score
Rank/total
AA-LCR
Thinking Mode
71
11 / 28

Claw-style Agent Evaluation

1 evaluations
Benchmark / mode
Score
Rank/total
Pinch Bench
Thinking ModeTools
88
6 / 38

Competitor Comparison

Benchmark scores for Claude Sonnet 4.6 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkClaude Sonnet 4.6CurrentClaude Opus 4.6GPT-5.2Gemini 3.0 Pro (Preview 11-2025)
ARC-AGI-2
Accuracy and cost per task
综合评估
58.30Thinking Enabled
66.30Extended Thinking
54.20Deep Thinking Mode
45.10Thinking Enabled
HLE
Accuracy
综合评估
49.00Thinking Enabled | Tools
53.00Extended Thinking | Tools
45.50Deep Thinking Mode | Tools
45.80Thinking Level · High | Tools
LiveBench
Accuracy
综合评估
75.47Thinking Level · Medium
76.33Thinking Level · High
74.84Thinking Level · High
73.39Thinking Level · High
GPQA Diamond
Accuracy
科学与综合推理
89.90Thinking Enabled
91.31Extended Thinking
93.20Deep Thinking Mode
93.80Thinking Enabled
SWE-bench Verified
Accuracy
编程与软件工程
79.60Thinking Enabled
80.84Extended Thinking | Tools
80.00Thinking Level · Extra High | Tools
76.20Thinking Enabled
Creative Writing
Elo、大模型评判两两对战
写作和创作
1804.50Standard Mode
1803.80Standard Mode
1698.90Standard Mode
--
FrontierMath - Tier 4
Accuracy
数学推理
8.3016K
22.90Thinking Level · High
18.80Thinking Level · Extra High
18.80Standard Mode
τ²-Bench - Telecom
Accuracy
Agent能力评测
97.90Thinking Enabled | Tools
99.25Extended Thinking | Tools
98.70Thinking Level · Extra High | Tools
98.00Thinking Level · High | Tools
BrowseComp
Accuracy
AI Agent - 信息收集
74.70Thinking Enabled | Tools
84.00Thinking Enabled | Tools
65.80Thinking Level · Extra High | Tools
59.20Thinking Level · High | Tools
MCP-Atlas
Pass rate / claim coverage
AI Agent - 工具使用
69.50Standard Mode | Tools
76.80Thinking Level · High | Tools
67.60Thinking Level · Extra High | Tools
70.30Standard Mode | Tools
OSWorld-Verified
Accuracy
AI Agent - 工具使用
72.50Thinking Enabled | Tools
72.70Extended Thinking | Tools
--
--
Terminal Bench 2.0
Accuracy
AI Agent - 工具使用
59.10Thinking Enabled | Tools
65.40Extended Thinking | Tools
--
56.90Thinking Level · High | Tools
4 additional benchmarks remain in the chart above.

Standard API Pricing: Claude Sonnet 4.6 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Claude Sonnet 4.6: Base price applies to <= 200K
Gemini 3.0 Pro (Preview 11-2025): Base price applies to <= 200000
ModelSupplierStandard inputStandard outputBase price applies to
Claude Sonnet 4.6
Anthropic$3 / 1M tokens$15 / 1M tokens<= 200K
Claude Opus 4.6
Anthropic$5 / 1M tokens$25 / 1M tokens
GPT-5.2
Facebook AI研究实验室$1.75 / 1M tokens$14 / 1M tokens
Gemini 3.0 Pro (Preview 11-2025)
Google Deep Mind$2 / 1M tokens$12 / 1M tokens<= 200000

Version History

How each version of the Claude Sonnet 4.6 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkClaude Sonnet 4.6CurrentClaude Sonnet 4.5Claude Sonnet 4Claude Sonnet 3.7
ARC-AGI-2
Accuracy and cost per task
综合评估
58.30Thinking Enabled
13.60Thinking Enabled
5.90Thinking Enabled
--
HLE
Accuracy
综合评估
49.00Thinking Enabled | Tools
33.60Thinking Enabled | Tools
9.60Thinking Enabled
10.30Thinking Enabled
LiveBench
Accuracy
综合评估
75.47Thinking Level · Medium
68.1964K
61.2764K
--
GPQA Diamond
Accuracy
科学与综合推理
89.90Thinking Enabled
83.40Thinking Enabled
83.80Deep Thinking Mode | Tools
77.00Thinking Enabled
SWE-bench Verified
Accuracy
编程与软件工程
79.60Thinking Enabled
82.00Thinking Enabled | Tools
80.20Thinking Enabled | Tools
70.30Thinking Enabled | Tools
Creative Writing
Elo、大模型评判两两对战
写作和创作
1804.50Standard Mode
1674.50Standard Mode
1480.30Standard Mode
--
FrontierMath - Tier 4
Accuracy
数学推理
8.3016K
4.2032K
0.00Standard Mode
--
τ²-Bench - Telecom
Accuracy
Agent能力评测
97.90Thinking Enabled | Tools
98.00Thinking Enabled | Tools
65.00Thinking Enabled | Tools
55.00Thinking Enabled | Tools
BrowseComp
Accuracy
AI Agent - 信息收集
74.70Thinking Enabled | Tools
24.10Thinking Enabled | Tools
--
--
MCP-Atlas
Pass rate / claim coverage
AI Agent - 工具使用
69.50Standard Mode | Tools
59.50Thinking Enabled | Tools
--
--
OSWorld-Verified
Accuracy
AI Agent - 工具使用
72.50Thinking Enabled | Tools
61.40Thinking Enabled | Tools
42.20Thinking Enabled | Tools
28.00Thinking Enabled | Tools
Terminal Bench 2.0
Accuracy
AI Agent - 工具使用
59.10Thinking Enabled | Tools
42.80Thinking Enabled | Tools
--
--
3 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: ARC-AGI-2 · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Claude Sonnet 4.6 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Claude Sonnet 4.6: Base price applies to <= 200K
Claude Sonnet 4.5: Base price applies to <= 200000
Claude Sonnet 4: Base price applies to <= 200000
ModelSupplierStandard inputStandard outputBase price applies to
Claude Sonnet 4.6
Anthropic$3 / 1M tokens$15 / 1M tokens<= 200K
Claude Sonnet 4.5
Anthropic$3 / 1M tokens$15 / 1M tokens<= 200000
Claude Sonnet 4
Anthropic$3 / 1M tokens$15 / 1M tokens<= 200000
Claude Sonnet 3.7
Anthropic$3 / 1M tokens$15 / 1M tokens

Sources