DataLearner logo

Claude Sonnet 4.6 Benchmark Details

Claude Sonnet 4.6 currently shows benchmark results led by Terminal Bench Hard (17 / 244, score 53), LiveBench (11 / 117, score 75.32), SWE-bench Verified (18 / 119, score 79.60). This page also compares it with 3 competitor models and 3 predecessor or same-series models, including performance and pricing views when available. 3 source links are attached for reference.

Benchmark Results

Claude Sonnet 4.6

Benchmark Results

Thinking
Tool usage
Internet

Abstract Generalization

5 evaluations
Benchmark / mode
Score
Rank/total
ARC-AGI-1
Thinking Level · HighTools
86.50
88 / 197
ARC-AGI-1
Thinking Level · MaxTools
86
91 / 197
ARC-AGI-2
Thinking Mode
58.33
80 / 185
ARC-AGI-2
Thinking Level · HighTools
60.42
74 / 185
ARC-AGI-2
Thinking Level · Max
58.33
80 / 185

Knowledge Exams

2 evaluations
Benchmark / mode
Score
Rank/total
HLE
Thinking Mode
33.20
109 / 235
HLE
Thinking ModeTools
49
46 / 235

Scientific Reasoning

6 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Level · Medium
83.33
104 / 255
GPQA Diamond
Thinking Mode
89.90
43 / 255
GPQA Diamond
Thinking Level · High
83.33
104 / 255
GPQA Diamond
Thinking Level · Max
78.79
141 / 255
CritPt
Standard Mode
0.90
162 / 204
CritPt
Thinking Level · Max
3.10
114 / 204

Repository Engineering

2 evaluations
Benchmark / mode
Score
Rank/total
SWE-bench Verified
Thinking Mode
79.60
18 / 119
DeepSWE
Thinking Level · HighTools
29.93
87 / 93

Writing

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1810.40
21 / 110

Mathematics

1 evaluations
Benchmark / mode
Score
Rank/total
8.30
34 / 80

Service Workflows

4 evaluations
Benchmark / mode
Score
Rank/total
τ²-Bench - Telecom
Thinking ModeTools
97.90
6 / 31
Public Benefits Bench v1.1
Thinking Level · MaxTools
62.45
15 / 26
SAGE
Thinking Level · MaxTools
46.58
27 / 64
τ³-Banking
Thinking Level · MaxTools
34.40
50 / 167

Fact Finding

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
Thinking ModeTools
74.70
30 / 58

Cross-capability Suites

3 evaluations
Benchmark / mode
Score
Rank/total
LiveBench
Thinking Level · Low
70.44
36 / 117
LiveBench
Thinking Level · Medium
72.99
23 / 117
LiveBench
Thinking Level · High
75.32
11 / 117

Agentic Development

4 evaluations
Benchmark / mode
Score
Rank/total
Terminal Bench 2.0
Thinking ModeTools
59.10
22 / 50
Terminal Bench Hard
Standard ModeTools
46.20
29 / 244
Terminal Bench Hard
Thinking Level · MaxTools
53
17 / 244
Terminal-Bench 4.0
Thinking Level · MaxTools
3
72 / 98

Memory & Persistence

4 evaluations
Benchmark / mode
Score
Rank/total
Context Arena
Standard Mode
55.18
80 / 126
Context Arena
Thinking Level · Low
82.80
30 / 126
Context Arena
Thinking Level · Medium
82.56
31 / 126
Context Arena
Thinking Level · High
82.82
29 / 126

Cross-industry Work

1 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA
Thinking Mode
57
5 / 15

Desktop Workflows

1 evaluations
Benchmark / mode
Score
Rank/total
OSWorld-Verified
Thinking ModeTools
72.50
19 / 28

Tool Orchestration

3 evaluations
Benchmark / mode
Score
Rank/total
Pinch Bench
Thinking ModeTools
88
6 / 38
MCP-Atlas
Standard ModeTools
69.50
32 / 44
62.68
31 / 45

Capability Indices

2 evaluations
Benchmark / mode
Score
Rank/total
ECI
unknown
152.25
37 / 167
Vals Index
Thinking Level · MaxTools
50.59
31 / 42

Visual Understanding

4 evaluations
Benchmark / mode
Score
Rank/total
MMMU-Pro
Standard Mode
70.60
125 / 229
MMMU-Pro
unknown
74.50
95 / 229
MMMU-Pro
Thinking ModeTools
75.60
82 / 229
MMMU-Pro
Thinking Level · Max
73.30
110 / 229

Scientific Computing

1 evaluations
Benchmark / mode
Score
Rank/total
SciCode
Thinking Level · Max
50.10
72 / 134

Legal

4 evaluations
Benchmark / mode
Score
Rank/total
Harvey Lab-AA
Standard ModeTools
84.44
23 / 44
Harvey Lab-AA
Thinking Level · MaxTools
85.97
19 / 44
Legal Research Bench
Thinking Level · MaxTools
38.46
24 / 45
Harvey's Legal Agent Benchmark
Thinking Level · MaxTools
5
23 / 43

Finance

2 evaluations
Benchmark / mode
Score
Rank/total
EMB (Excel Modeling)
Thinking Level · MaxTools
60.15
28 / 45
Finance Agent v2
Thinking Level · MaxTools
51.03
28 / 43

Documents & Charts

1 evaluations
Benchmark / mode
Score
Rank/total
GDP.pdf
Thinking Level · Max
15.80
57 / 122

Data Analysis

1 evaluations
Benchmark / mode
Score
Rank/total
AA-AnalystAgent
Thinking Level · MaxToolsInternet
20
18 / 29

Maintenance & Optimization

1 evaluations
Benchmark / mode
Score
Rank/total
Code Migration
Thinking Level · MaxTools
39.89
24 / 47

Code Generation & Editing

1 evaluations
Benchmark / mode
Score
Rank/total
Vibe Code Bench v1.1
Thinking Level · MaxTools
51.48
39 / 64

Competitor Comparison

Benchmark scores for Claude Sonnet 4.6 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkClaude Sonnet 4.6CurrentClaude Opus 4.6GPT-5.2Gemini 3.0 Pro (Preview 11-2025)
ARC-AGI-1
Score (%)
Abstract Generalization
86.50Thinking Level · High | Tools
94.00Thinking Level · High
90.50Deep Thinking Mode
87.50Thinking Enabled
ARC-AGI-2
Score (% solved); cost per task (USD)
Abstract Generalization
60.42Thinking Level · High | Tools
69.17Thinking Level · High
54.20Deep Thinking Mode
45.10Thinking Enabled
HLE
Accuracy
Knowledge Exams
49.00Thinking Enabled | Tools
53.00Extended Thinking | Tools
45.50Deep Thinking Mode | Tools
45.80Thinking Level · High | Tools
CritPt
Score
Scientific Reasoning
3.10Thinking Level · High
12.60Thinking Level · High
11.60Thinking Level · Extra High
9.10Thinking Level · High
GPQA Diamond
Accuracy
Scientific Reasoning
89.90Thinking Enabled
91.31Extended Thinking
93.20Deep Thinking Mode
93.80Thinking Enabled
SWE-bench Verified
Accuracy
Repository Engineering
79.60Thinking Enabled
80.84Extended Thinking | Tools
80.00Thinking Level · Extra High | Tools
76.20Thinking Enabled
Creative Writing
Elo、大模型评判两两对战
Writing
1810.40Standard Mode
1809.10Standard Mode
1702.90Standard Mode
--
FrontierMath - Tier 4
Accuracy
Mathematics
8.3016K
22.90Thinking Level · High
18.80Thinking Level · Extra High
18.80Standard Mode
SAGE
Accuracy (%)
Service Workflows
46.58Thinking Level · High | Tools
51.58Thinking Level · High | Tools
49.27Thinking Level · Extra High | Tools
--
τ²-Bench - Telecom
Accuracy
Service Workflows
97.90Thinking Enabled | Tools
99.25Extended Thinking | Tools
89.69Thinking Level · High | Tools
98.00Thinking Enabled | Tools
τ³-Banking
Score
Service Workflows
34.40Thinking Level · High | Tools
27.32Thinking Level · High | Tools
32.22Thinking Level · High | Tools
--
BrowseComp
Accuracy
Fact Finding
74.70Thinking Enabled | Tools
84.00Thinking Enabled | Tools
65.80Thinking Level · Extra High | Tools
59.20Thinking Level · High | Tools
12 additional benchmarks remain in the chart above.

Standard API Pricing: Claude Sonnet 4.6 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Gemini 3.0 Pro (Preview 11-2025): Base price applies to <= 200000
ModelSupplierStandard inputStandard outputBase price applies to
Claude Sonnet 4.6
Anthropic$3 / 1M tokens$15 / 1M tokens—
Claude Opus 4.6
Anthropic$5 / 1M tokens$25 / 1M tokens—
GPT-5.2
OpenAI$1.75 / 1M tokens$14 / 1M tokens—
Gemini 3.0 Pro (Preview 11-2025)
Google DeepMind$2 / 1M tokens$12 / 1M tokens<= 200000

Version History

How each version of the Claude Sonnet 4.6 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkClaude Sonnet 4.6CurrentClaude Sonnet 4.5Claude Sonnet 4Claude Sonnet 3.7
ARC-AGI-1
Score (%)
Abstract Generalization
86.50Thinking Level · High | Tools
63.67Thinking Enabled
40.00Thinking Enabled
--
ARC-AGI-2
Score (% solved); cost per task (USD)
Abstract Generalization
60.42Thinking Level · High | Tools
13.61Thinking Enabled
5.93Thinking Enabled
--
HLE
Accuracy
Knowledge Exams
49.00Thinking Enabled | Tools
33.60Thinking Enabled | Tools
9.60Thinking Enabled
10.30Thinking Enabled
CritPt
Score
Scientific Reasoning
3.10Thinking Level · High
1.10Thinking Enabled
1.10Standard Mode
0.90Thinking Enabled
GPQA Diamond
Accuracy
Scientific Reasoning
89.90Thinking Enabled
73.70Standard Mode
83.80Deep Thinking Mode | Tools
77.00Thinking Enabled
SWE-bench Verified
Accuracy
Repository Engineering
79.60Thinking Enabled
82.00Thinking Enabled | Tools
80.20Thinking Enabled | Tools
70.30Thinking Enabled | Tools
Creative Writing
Elo、大模型评判两两对战
Writing
1810.40Standard Mode
1677.60Standard Mode
1482.80Standard Mode
1411.70Standard Mode
FrontierMath - Tier 4
Accuracy
Mathematics
8.3016K
4.2032K
0.00Standard Mode
--
SAGE
Accuracy (%)
Service Workflows
46.58Thinking Level · High | Tools
36.06Thinking Enabled | Tools
35.00Standard Mode | Tools
--
τ²-Bench - Telecom
Accuracy
Service Workflows
97.90Thinking Enabled | Tools
98.00Thinking Enabled | Tools
65.00Thinking Enabled | Tools
55.00Thinking Enabled | Tools
τ³-Banking
Score
Service Workflows
34.40Thinking Level · High | Tools
24.50Thinking Enabled | Tools
16.70Thinking Enabled | Tools
--
BrowseComp
Accuracy
Fact Finding
74.70Thinking Enabled | Tools
24.10Thinking Enabled | Tools
--
--
13 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: ARC-AGI-1 · Abstract Generalization

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Claude Sonnet 4.6 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Claude Sonnet 4.5: Base price applies to <= 200000
Claude Sonnet 4: Base price applies to <= 200000
ModelSupplierStandard inputStandard outputBase price applies to
Claude Sonnet 4.6
Anthropic$3 / 1M tokens$15 / 1M tokens—
Claude Sonnet 4.5
Anthropic$3 / 1M tokens$15 / 1M tokens<= 200000
Claude Sonnet 4
Anthropic$3 / 1M tokens$15 / 1M tokens<= 200000
Claude Sonnet 3.7
Anthropic$3 / 1M tokens$15 / 1M tokens—

Sources