DataLearner logo

Claude Sonnet 4.5 Benchmark Details

Claude Sonnet 4.5 currently shows benchmark results led by AIME2025 (1 / 107, score 100), MMLU Pro (7 / 132, score 88), SWE-bench Verified (8 / 113, score 82). This page also compares it with 2 competitor models and 4 predecessor or same-series models, including performance and pricing views when available. 2 source links are attached for reference.

Benchmark Results

Claude Sonnet 4.5

Benchmark Results

Thinking
Tool usage

General Knowledge

12 evaluations
Benchmark / mode
Score
Rank/total
88
7 / 132
83.40
64 / 188
73.70
104 / 188
LiveBench
Standard Mode
53.69
83 / 115
68.19
46 / 115
63.70
35 / 68
25.50
55 / 68
33.60
80 / 173
17.70
127 / 173
7.10
160 / 173
13.60
38 / 62
3.80
52 / 62

Coding and Software Engineer

6 evaluations
Benchmark / mode
Score
Rank/total
CodeClash
Standard ModeTools
1389
1 / 8
77.20
28 / 113
71
48 / 123
59
72 / 123

Math and Reasoning

8 evaluations
Benchmark / mode
Score
Rank/total
100
1 / 107
87
46 / 107
37
97 / 107
27.10
8 / 16
5.20
38 / 60
2.10
56 / 80
4.20
40 / 80

AI Agent - Tool Usage

5 evaluations
Benchmark / mode
Score
Rank/total
61.40
21 / 25
MCP-Atlas
Thinking EnabledTools
59.50
22 / 28
42.80
42 / 47

Multimodal Understanding

1 evaluations
Benchmark / mode
Score
Rank/total
77.80
15 / 29

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
Simple Bench
Standard Mode
54.30
22 / 63

Agent Level Benchmark

4 evaluations
Benchmark / mode
Score
Rank/total

Instruction Following

1 evaluations
Benchmark / mode
Score
Rank/total
57.30
23 / 31

AI Agent - Information Search

1 evaluations
Benchmark / mode
Score
Rank/total
24.10
51 / 53

Productivity Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
39
16 / 21

Long Context

1 evaluations
Benchmark / mode
Score
Rank/total
66
11 / 16

Claw-style Agent Evaluation

2 evaluations
Benchmark / mode
Score
Rank/total
Pinch Bench
Thinking EnabledTools
88.20
4 / 37
Claw Bench
Thinking EnabledTools
88.10
13 / 29

Competitor Comparison

Benchmark scores for Claude Sonnet 4.5 compared against top models in its class

Claude Sonnet 4.5GPT-5.1Gemini 2.5-Pro
Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkClaude Sonnet 4.5CurrentGPT-5.1Gemini 2.5-Pro
ARC-AGI
综合评估
63.70Thinking Enabled
72.80Thinking Level · High
37.00Thinking Enabled
ARC-AGI-2
综合评估
13.60Thinking Enabled
17.60Thinking Level · High
4.90Thinking Enabled
GPQA Diamond
综合评估
83.40Thinking Enabled
88.10Thinking Enabled
86.40Thinking Enabled
HLE
综合评估
33.60Thinking Enabled | Tools
42.70Thinking Level · High | Tools
21.60Thinking Enabled
LiveBench
综合评估
68.1964K
72.04Thinking Level · High
58.33Thinking Level · High
MMLU Pro
综合评估
88.00Thinking Enabled
--
86.00Standard Mode
CodeClash
编程与软件工程
1389.00Standard Mode | Tools
--
1125.00Standard Mode | Tools
LiveCodeBench
编程与软件工程
71.00Thinking Enabled
--
77.10Standard Mode
SWE-Bench Pro - Public
编程与软件工程
43.60Thinking Enabled
50.80Thinking Level · High
--
SWE-bench Verified
编程与软件工程
82.00Thinking Enabled | Tools
76.30Thinking Level · High
67.20Thinking Enabled
AIME2025
数学推理
100.00Thinking Enabled | Tools
94.00Thinking Level · High
88.00Thinking Enabled
FrontierMath
数学推理
5.20Standard Mode
26.70Thinking Level · High | Tools
11.00Standard Mode
14 additional benchmarks remain in the chart above.

Standard API Pricing: Claude Sonnet 4.5 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Claude Sonnet 4.5: Base price applies to <= 200000
Gemini 2.5-Pro: Base price applies to <= 200000
ModelSupplierStandard inputStandard outputBase price applies to
Claude Sonnet 4.5
Anthropic$3 / 1M tokens$15 / 1M tokens<= 200000
GPT-5.1
OpenAI$1.25 / 1M tokens$10 / 1M tokens
Gemini 2.5-Pro
Google Deep Mind$1.25 / 1M tokens$10 / 1M tokens<= 200000

Version History

How each version of the Claude Sonnet 4.5 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkClaude Sonnet 4.5CurrentClaude Sonnet 4Claude Sonnet 3.7Claude 3.5 Sonnet NewClaude 3.5 Sonnet
ARC-AGI
综合评估
63.70Thinking Enabled
40.00Thinking Enabled
--
--
--
ARC-AGI-2
综合评估
13.60Thinking Enabled
5.90Thinking Enabled
--
--
--
GPQA Diamond
综合评估
83.40Thinking Enabled
83.80Deep Thinking Mode | Tools
77.00Thinking Enabled
65.00Standard Mode
59.40Standard Mode
HLE
综合评估
33.60Thinking Enabled | Tools
9.60Thinking Enabled
10.30Thinking Enabled
--
--
LiveBench
综合评估
68.1964K
61.2764K
--
--
--
MMLU Pro
综合评估
88.00Thinking Enabled
84.00Thinking Enabled
--
78.00Standard Mode
77.64Standard Mode
CodeClash
编程与软件工程
1389.00Standard Mode | Tools
1223.00Standard Mode | Tools
--
--
--
LiveCodeBench
编程与软件工程
71.00Thinking Enabled
66.00Thinking Enabled
--
38.70Standard Mode
--
SWE-Bench Pro - Public
编程与软件工程
43.60Thinking Enabled
42.70Thinking Enabled
--
--
--
SWE-bench Verified
编程与软件工程
82.00Thinking Enabled | Tools
80.20Thinking Enabled | Tools
70.30Thinking Enabled | Tools
49.00Standard Mode
--
AIME2025
数学推理
100.00Thinking Enabled | Tools
85.00Deep Thinking Mode | Tools
54.80Standard Mode
--
--
FrontierMath
数学推理
5.20Standard Mode
4.10Standard Mode
4.10Thinking Enabled
2.10Standard Mode
1.00Standard Mode
14 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: ARC-AGI · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Claude Sonnet 4.5 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Claude Sonnet 4.5: Base price applies to <= 200000
Claude Sonnet 4: Base price applies to <= 200000
ModelSupplierStandard inputStandard outputBase price applies to
Claude Sonnet 4.5
Anthropic$3 / 1M tokens$15 / 1M tokens<= 200000
Claude Sonnet 4
Anthropic$3 / 1M tokens$15 / 1M tokens<= 200000
Claude Sonnet 3.7
Anthropic$3 / 1M tokens$15 / 1M tokens
Claude 3.5 Sonnet New
Anthropic$3 / 1M tokens$15 / 1M tokens
Claude 3.5 Sonnet
Anthropic$3 / 1M tokens$15 / 1M tokens

Sources