DataLearner logo

Claude Sonnet 4.5 Benchmark Details

Claude Sonnet 4.5 currently shows benchmark results led by AIME2025 (1 / 215, score 100), τ²-Bench - Telecom (8 / 264, score 98), MMLU-Pro (8 / 175, score 88). This page also compares it with 2 competitor models and 4 predecessor or same-series models, including performance and pricing views when available. 2 source links are attached for reference.

Benchmark Results

Claude Sonnet 4.5

Benchmark Results

Thinking
Tool usage
Parallel

Knowledge Exams

6 evaluations
Benchmark / mode
Score
Rank/total
MMLU-Pro
Thinking Enabled
88
8 / 175
HLE
Standard Mode
7.20
442 / 568
HLE
Standard Mode
7.10
445 / 568
HLE
Thinking Enabled
17.80
319 / 568
HLE
Thinking Enabled
17.70
320 / 568
HLE
Thinking EnabledTools
33.60
186 / 568

Abstract Generalization

4 evaluations
Benchmark / mode
Score
Rank/total
ARC-AGI-1
Standard Mode
25.50
154 / 176
ARC-AGI-1
Thinking Enabled
63.67
109 / 176
ARC-AGI-2
Standard Mode
3.80
137 / 164
ARC-AGI-2
Thinking Enabled
13.61
108 / 164

Scientific Reasoning

3 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
73.70
293 / 461
GPQA Diamond
Thinking Enabled
83.40
186 / 461
CritPt
Thinking Enabled
1.10
151 / 204

Repository Engineering

3 evaluations
Benchmark / mode
Score
Rank/total
SWE-bench Verified
Thinking EnabledTools
77.20
28 / 116
SWE-bench Verified
Thinking Level · HighTools
82
8 / 116
SWE-Bench Pro - Public
Thinking Enabled
43.60
54 / 62

Algorithmic Coding

3 evaluations
Benchmark / mode
Score
Rank/total
CodeClash
Standard ModeTools
1389
1 / 8
LiveCodeBench
Standard Mode
59
137 / 251
LiveCodeBench
Thinking Enabled
71
85 / 251

Mathematics

10 evaluations
Benchmark / mode
Score
Rank/total
AIME2025
Standard Mode
37
174 / 215
AIME2025
Thinking Enabled
87
69 / 215
AIME2025
Thinking EnabledTools
100
1 / 215
IMO-ProofBench
Thinking Enabled
27.10
8 / 16
23.86
48 / 58
FrontierMath
Standard Mode
5.20
38 / 60
IMO-ProofBench Advanced
Thinking Enabled
4.80
19 / 24
2.10
56 / 80
4.20
40 / 80

Writing

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1674.40
30 / 106

Agentic Development

6 evaluations
Benchmark / mode
Score
Rank/total
Terminal-Bench 2.1
Thinking EnabledTools
55.80
127 / 199
Terminal-Bench
Standard ModeTools
27
25 / 35
Terminal-Bench
Thinking EnabledTools
50
3 / 35
Terminal Bench 2.0
Thinking EnabledTools
42.80
43 / 48
Terminal Bench Hard
Standard ModeTools
28.80
116 / 244
Terminal Bench Hard
Thinking EnabledTools
33
92 / 244

Visual Understanding

7 evaluations
Benchmark / mode
Score
Rank/total
MMMU
unknown
77.80
23 / 74
MMMU
Thinking Enabled
77.80
23 / 74
MMMU-Pro
Standard Mode
65.20
153 / 229
MMMU-Pro
unknown
68.90
140 / 229
MMMU-Pro
Thinking Enabled
68.70
143 / 229
VPCT
Standard Mode
38
20 / 24
VPCT
32K
39.80
17 / 24

Commonsense

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Standard Mode
54.30
47 / 96

Service Workflows

7 evaluations
Benchmark / mode
Score
Rank/total
τ²-Bench - Telecom
Standard ModeTools
70.50
138 / 264
τ²-Bench - Telecom
Thinking EnabledTools
98
8 / 264
τ²-Bench
Standard ModeTools
71
26 / 44
τ²-Bench
Thinking EnabledTools
84.70
9 / 44
SAGE
Standard ModeTools
32.88
51 / 64
SAGE
Thinking EnabledTools
36.06
47 / 64
τ³-Banking
Thinking EnabledTools
24.50
81 / 167

Instruction Following

3 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Standard Mode
42.70
199 / 282
IF Bench
Thinking Enabled
57.30
127 / 282
IF Bench
Thinking EnabledTools
57.30
127 / 282

Fact Finding

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
Thinking EnabledTools
24.10
56 / 58

Cross-capability Suites

2 evaluations
Benchmark / mode
Score
Rank/total
LiveBench
Standard Mode
53.69
85 / 117
68.19
47 / 117

Cross-industry Work

1 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA
Thinking Enabled
39
10 / 15

Long Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
AA-LCR
Standard Mode
54
137 / 174
AA-LCR
Thinking Enabled
72.30
96 / 174

Desktop Workflows

1 evaluations
Benchmark / mode
Score
Rank/total
OSWorld-Verified
Thinking EnabledTools
61.40
24 / 28

Tool Orchestration

3 evaluations
Benchmark / mode
Score
Rank/total
Pinch Bench
Thinking EnabledTools
88.20
5 / 38
Claw Bench
Thinking EnabledTools
88.10
13 / 29
MCP-Atlas
Thinking EnabledTools
59.50
37 / 44

Scientific Computing

1 evaluations
Benchmark / mode
Score
Rank/total
SciCode
Thinking Enabled
45.70
90 / 134

Maintenance & Optimization

1 evaluations
Benchmark / mode
Score
Rank/total
GSO
Standard ModeTools
14.70
12 / 21

ML Engineering

2 evaluations
Benchmark / mode
Score
Rank/total
WeirdML v2
Standard ModeTools
46.70
40 / 52
WeirdML v2
16KTools
47.71
39 / 52

Preference Arenas

2 evaluations
Benchmark / mode
Score
Rank/total
Text Arena (Coding)
Standard Mode
1387
32 / 35
1391
31 / 35

Capability Frontier Metrics

1 evaluations
Benchmark / mode
Score
Rank/total
121.95
12 / 22

Documents & Charts

1 evaluations
Benchmark / mode
Score
Rank/total
GDP.pdf
Thinking Enabled
5.20
99 / 122

Code Generation & Editing

1 evaluations
Benchmark / mode
Score
Rank/total
Vibe Code Bench v1.1
Thinking EnabledTools
22.62
45 / 60

Clinical Workflows

4 evaluations
Benchmark / mode
Score
Rank/total
MedScribe
Standard ModeTools
84.52
22 / 66
MedScribe
Thinking EnabledTools
84.10
26 / 66
MedCode
Standard ModeTools
40.57
45 / 64
MedCode
Thinking EnabledTools
44.13
31 / 64

Capability Indices

1 evaluations
Benchmark / mode
Score
Rank/total
ECI
unknown
146.84
61 / 167

Memory & Persistence

1 evaluations
Benchmark / mode
Score
Rank/total
EBR-bench
Standard ModeTools
2.38
22 / 23

Competitor Comparison

Benchmark scores for Claude Sonnet 4.5 compared against top models in its class

Claude Sonnet 4.5GPT-5.1Gemini 2.5-Pro
Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkClaude Sonnet 4.5CurrentGPT-5.1Gemini 2.5-Pro
HLE
Accuracy
Knowledge Exams
33.60Thinking Enabled | Tools
42.70Thinking Level · High | Tools
22.50Thinking Enabled
MMLU-Pro
Accuracy
Knowledge Exams
88.00Thinking Enabled
--
86.00Standard Mode
ARC-AGI-1
Score (%)
Abstract Generalization
63.67Thinking Enabled
72.83Thinking Level · High
37.00Thinking Enabled
ARC-AGI-2
Score (% solved); cost per task (USD)
Abstract Generalization
13.61Thinking Enabled
17.64Thinking Level · High
4.86Thinking Enabled
CritPt
Score
Scientific Reasoning
1.10Thinking Enabled
4.90Thinking Level · High
2.60Thinking Enabled
GPQA Diamond
Accuracy
Scientific Reasoning
83.40Thinking Enabled
88.10Thinking Enabled
86.40Thinking Enabled
SWE-Bench Pro - Public
Accuracy
Repository Engineering
43.60Thinking Enabled
50.80Thinking Level · High
--
SWE-bench Verified
Accuracy
Repository Engineering
82.00Thinking Enabled | Tools
76.30Thinking Level · High | Tools
67.20Thinking Enabled
CodeClash
Elo / win rate
Algorithmic Coding
1389.00Standard Mode | Tools
--
1125.00Standard Mode | Tools
LiveCodeBench
Pass @K
Algorithmic Coding
71.00Thinking Enabled
86.80Thinking Level · High
80.10Thinking Enabled
AIME2025
Accuracy
Mathematics
100.00Thinking Enabled | Tools
94.17Thinking Level · High | Tools
88.00Thinking Enabled
FrontierMath
Accuracy
Mathematics
5.20Standard Mode
26.70Thinking Level · High | Tools
11.00Standard Mode
27 additional benchmarks remain in the chart above.

Standard API Pricing: Claude Sonnet 4.5 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Claude Sonnet 4.5: Base price applies to <= 200000
Gemini 2.5-Pro: Base price applies to <= 200000
ModelSupplierStandard inputStandard outputBase price applies to
Claude Sonnet 4.5
Anthropic$3 / 1M tokens$15 / 1M tokens<= 200000
GPT-5.1
OpenAI$1.25 / 1M tokens$10 / 1M tokens—
Gemini 2.5-Pro
Google DeepMind$1.25 / 1M tokens$10 / 1M tokens<= 200000

Version History

How each version of the Claude Sonnet 4.5 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkClaude Sonnet 4.5CurrentClaude Sonnet 4Claude Sonnet 3.7Claude 3.5 Sonnet NewClaude 3.5 Sonnet
HLE
Accuracy
Knowledge Exams
33.60Thinking Enabled | Tools
10.70Thinking Enabled
10.30Thinking Enabled
--
4.08Thinking Level · High
MMLU-Pro
Accuracy
Knowledge Exams
88.00Thinking Enabled
84.00Thinking Enabled
--
78.00Standard Mode
77.64Thinking Level · High
ARC-AGI-1
Score (%)
Abstract Generalization
63.67Thinking Enabled
40.00Thinking Enabled
--
--
--
ARC-AGI-2
Score (% solved); cost per task (USD)
Abstract Generalization
13.61Thinking Enabled
5.93Thinking Enabled
--
--
--
CritPt
Score
Scientific Reasoning
1.10Thinking Enabled
1.10Standard Mode
0.90Thinking Enabled
--
--
GPQA Diamond
Accuracy
Scientific Reasoning
83.40Thinking Enabled
83.80Deep Thinking Mode | Tools
77.00Thinking Enabled
65.00Standard Mode
59.40Standard Mode
SWE-Bench Pro - Public
Accuracy
Repository Engineering
43.60Thinking Enabled
42.70Thinking Enabled
--
--
--
SWE-bench Verified
Accuracy
Repository Engineering
82.00Thinking Enabled | Tools
80.20Thinking Enabled | Tools
70.30Thinking Enabled | Tools
49.00Standard Mode
--
CodeClash
Elo / win rate
Algorithmic Coding
1389.00Standard Mode | Tools
1223.00Standard Mode | Tools
--
--
--
LiveCodeBench
Pass @K
Algorithmic Coding
71.00Thinking Enabled
66.00Thinking Enabled
47.30Thinking Enabled
38.70Standard Mode
38.10Standard Mode
AIME2025
Accuracy
Mathematics
100.00Thinking Enabled | Tools
85.00Deep Thinking Mode | Tools
56.30Thinking Enabled
--
--
FrontierMath
Accuracy
Mathematics
5.20Standard Mode
4.10Standard Mode
4.10Thinking Enabled
2.10Standard Mode
1.00Standard Mode
27 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: HLE · Knowledge Exams

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Claude Sonnet 4.5 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Claude Sonnet 4.5: Base price applies to <= 200000
Claude Sonnet 4: Base price applies to <= 200000
ModelSupplierStandard inputStandard outputBase price applies to
Claude Sonnet 4.5
Anthropic$3 / 1M tokens$15 / 1M tokens<= 200000
Claude Sonnet 4
Anthropic$3 / 1M tokens$15 / 1M tokens<= 200000
Claude Sonnet 3.7
Anthropic$3 / 1M tokens$15 / 1M tokens—
Claude 3.5 Sonnet New
Anthropic$3 / 1M tokens$15 / 1M tokens—
Claude 3.5 Sonnet
Anthropic$3 / 1M tokens$15 / 1M tokens—

Sources