DataLearner logo

Claude Sonnet 4.6 Benchmark Details

Claude Sonnet 4.6 currently shows benchmark results led by LiveBench (12 / 115, score 75.47), Pinch Bench (6 / 38, score 88), SWE-bench Verified (18 / 114, score 79.60). This page also compares it with 3 competitor models and 3 predecessor or same-series models, including performance and pricing views when available. 3 source links are attached for reference.

Benchmark Results

Claude Sonnet 4.6

Benchmark Results

Thinking
Tool usage

General Knowledge

6 evaluations
Benchmark / mode
Score
Rank/total
LiveBench
Thinking Level · Low
70.44
36 / 115
LiveBench
Thinking Level · Medium
75.47
12 / 115
LiveBench
Thinking Level · High
75.32
15 / 115
58.30
21 / 62
49
34 / 181
33.20
86 / 181

Other

1 evaluations
Benchmark / mode
Score
Rank/total
89.90
42 / 226

Coding and Software Engineer

2 evaluations
Benchmark / mode
Score
Rank/total
79.60
18 / 114
DeepSWE
Thinking Level · HighTools
30
24 / 26

Math and Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
8.30
34 / 80

Agent Level Benchmark

1 evaluations
Benchmark / mode
Score
Rank/total

AI Agent - Information Search

1 evaluations
Benchmark / mode
Score
Rank/total
74.70
27 / 54

AI Agent - Tool Usage

3 evaluations
Benchmark / mode
Score
Rank/total
72.50
17 / 26
MCP-Atlas
Standard ModeTools
69.50
27 / 38
59.10
22 / 48

Productivity Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
57
11 / 21

Long Context

1 evaluations
Benchmark / mode
Score
Rank/total
71
4 / 18

Claw-style Agent Evaluation

1 evaluations
Benchmark / mode
Score
Rank/total
Pinch Bench
Thinking EnabledTools
88
6 / 38

Competitor Comparison

Benchmark scores for Claude Sonnet 4.6 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkClaude Sonnet 4.6CurrentClaude Opus 4.6GPT-5.2Gemini 3.0 Pro (Preview 11-2025)
ARC-AGI-2
综合评估
58.30Thinking Enabled
66.30Extended Thinking
54.20Deep Thinking Mode
45.10Thinking Enabled
HLE
综合评估
49.00Thinking Enabled | Tools
53.00Extended Thinking | Tools
45.50Deep Thinking Mode | Tools
45.80Thinking Level · High | Tools
LiveBench
综合评估
75.47Thinking Level · Medium
76.33Thinking Level · High
74.84Thinking Level · High
73.39Thinking Level · High
GPQA Diamond
科学与综合推理
89.90Thinking Enabled
91.31Extended Thinking
93.20Deep Thinking Mode
93.80Thinking Enabled
SWE-bench Verified
编程与软件工程
79.60Thinking Enabled
80.84Extended Thinking | Tools
80.00Thinking Level · Extra High | Tools
76.20Thinking Enabled
8.3016K
22.90Thinking Level · High
18.80Thinking Level · Extra High
18.80Thinking Enabled
τ²-Bench - Telecom
Agent能力评测
97.90Thinking Enabled | Tools
99.25Extended Thinking | Tools
98.70Thinking Level · Extra High | Tools
98.00Thinking Level · High | Tools
BrowseComp
AI Agent - 信息收集
74.70Thinking Enabled | Tools
84.00Thinking Enabled | Tools
65.80Thinking Level · Extra High | Tools
59.20Thinking Level · High | Tools
MCP-Atlas
AI Agent - 工具使用
69.50Standard Mode | Tools
76.80Thinking Level · High | Tools
67.60Thinking Level · Extra High | Tools
70.30Standard Mode | Tools
OSWorld-Verified
AI Agent - 工具使用
72.50Thinking Enabled | Tools
72.70Extended Thinking | Tools
--
--
Terminal Bench 2.0
AI Agent - 工具使用
59.10Thinking Enabled | Tools
65.40Extended Thinking | Tools
--
56.90Thinking Level · High | Tools
GDPval-AA
生产力知识
57.00Thinking Enabled
1606.00Extended Thinking | Tools
70.90Thinking Level · High | Tools
35.00Thinking Level · High
2 additional benchmarks remain in the chart above.

Standard API Pricing: Claude Sonnet 4.6 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Claude Sonnet 4.6: Base price applies to <= 200K
Claude Opus 4.6: Base price applies to <= 200K
Gemini 3.0 Pro (Preview 11-2025): Base price applies to <= 200000
ModelSupplierStandard inputStandard outputBase price applies to
Claude Sonnet 4.6
Anthropic$3 / 1M tokens$15 / 1M tokens<= 200K
Claude Opus 4.6
Anthropic$5 / 1M tokens$25 / 1M tokens<= 200K
GPT-5.2
Facebook AI研究实验室$1.75 / 1M tokens$14 / 1M tokens
Gemini 3.0 Pro (Preview 11-2025)
Google Deep Mind$2 / 1M tokens$12 / 1M tokens<= 200000

Version History

How each version of the Claude Sonnet 4.6 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkClaude Sonnet 4.6CurrentClaude Sonnet 4.5Claude Sonnet 4Claude Sonnet 3.7
ARC-AGI-2
综合评估
58.30Thinking Enabled
13.60Thinking Enabled
5.90Thinking Enabled
--
HLE
综合评估
49.00Thinking Enabled | Tools
33.60Thinking Enabled | Tools
9.60Thinking Enabled
10.30Thinking Enabled
LiveBench
综合评估
75.47Thinking Level · Medium
68.1964K
61.2764K
--
GPQA Diamond
科学与综合推理
89.90Thinking Enabled
83.40Thinking Enabled
83.80Deep Thinking Mode | Tools
77.00Thinking Enabled
SWE-bench Verified
编程与软件工程
79.60Thinking Enabled
82.00Thinking Enabled | Tools
80.20Thinking Enabled | Tools
70.30Thinking Enabled | Tools
8.3016K
4.2032K
0.00Standard Mode
--
τ²-Bench - Telecom
Agent能力评测
97.90Thinking Enabled | Tools
98.00Thinking Enabled | Tools
65.00Thinking Enabled | Tools
55.00Thinking Enabled | Tools
BrowseComp
AI Agent - 信息收集
74.70Thinking Enabled | Tools
24.10Thinking Enabled | Tools
--
--
MCP-Atlas
AI Agent - 工具使用
69.50Standard Mode | Tools
59.50Thinking Enabled | Tools
--
--
OSWorld-Verified
AI Agent - 工具使用
72.50Thinking Enabled | Tools
61.40Thinking Enabled | Tools
42.20Thinking Enabled | Tools
28.00Thinking Enabled | Tools
Terminal Bench 2.0
AI Agent - 工具使用
59.10Thinking Enabled | Tools
42.80Thinking Enabled | Tools
--
--
GDPval-AA
生产力知识
57.00Thinking Enabled
39.00Thinking Enabled
33.00Thinking Enabled
28.00Thinking Enabled
2 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: ARC-AGI-2 · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Claude Sonnet 4.6 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Claude Sonnet 4.6: Base price applies to <= 200K
Claude Sonnet 4.5: Base price applies to <= 200000
Claude Sonnet 4: Base price applies to <= 200000
ModelSupplierStandard inputStandard outputBase price applies to
Claude Sonnet 4.6
Anthropic$3 / 1M tokens$15 / 1M tokens<= 200K
Claude Sonnet 4.5
Anthropic$3 / 1M tokens$15 / 1M tokens<= 200000
Claude Sonnet 4
Anthropic$3 / 1M tokens$15 / 1M tokens<= 200000
Claude Sonnet 3.7
Anthropic$3 / 1M tokens$15 / 1M tokens

Sources