DataLearner logo

Claude Sonnet 4.5vsClaude Sonnet 4

Across 26 shared benchmarks, Claude Sonnet 4.5 leads overall: Claude Sonnet 4.5 wins 22, Claude Sonnet 4 wins 2, with 2 ties and an average score difference of +19.53.

Anthropic
Claude Sonnet 4.5

Anthropic · 2025-09-30 · Chat model

Anthropic
Claude Sonnet 4

Anthropic · 2025-05-23 · Reasoning model

Claude Sonnet 4.522 wins(85%)Ties2(8%)2 winsClaude Sonnet 4

Benchmark scores

Grouped by capability, sorted by largest gap within each. 26 shared benchmarks.

Agent Level Benchmark

Claude Sonnet 4.5 4/4
BenchmarkClaude Sonnet 4.5Claude Sonnet 4Diff
τ²-Bench7126 / 44Normal (With Tools)5235 / 44Normal (With Tools)+19
τ²-Bench - Telecom70.50138 / 264Normal (With Tools)52.30169 / 264Normal (With Tools)+18.20
τ³-Banking24.5080 / 164Thinking (With Tools)16.70100 / 164Thinking (With Tools)+7.80
Terminal Bench Hard28.80116 / 244Normal (With Tools)27.30120 / 244Normal (With Tools)+1.50

Coding and Software Engineer

Claude Sonnet 4.5 4/4
BenchmarkClaude Sonnet 4.5Claude Sonnet 4Diff
CodeClash1,3891 / 8Normal (With Tools)1,2234 / 8Normal (With Tools)+166
LiveCodeBench59136 / 250Normal (No Tools)48.50170 / 250Normal (No Tools)+10.50
SWE-bench Verified828 / 116Parallel Thinking (With Tools)80.2014 / 116Parallel Thinking (With Tools)+1.80
SWE-Bench Pro - Public43.6054 / 62Thinking (No Tools)42.7055 / 62Thinking (No Tools)+0.90

Math and Reasoning

Even 4/4
BenchmarkClaude Sonnet 4.5Claude Sonnet 4Diff
FrontierMath5.2038 / 60Normal (No Tools)4.1041 / 60Normal (No Tools)+1.10
AIME202537173 / 215Normal (No Tools)38170 / 215Normal (No Tools)-1
IMO-ProofBench27.108 / 16Thinking (No Tools)27.108 / 16Thinking (No Tools)
IMO-ProofBench Advanced4.8019 / 24Thinking (No Tools)4.8019 / 24Thinking (No Tools)

AI Agent - Tool Usage

Claude Sonnet 4.5 3/3
BenchmarkClaude Sonnet 4.5Claude Sonnet 4Diff
Terminal-Bench 2.155.80120 / 192Thinking (With Tools)36.30150 / 192Thinking (With Tools)+19.50
OSWorld-Verified61.4022 / 26Thinking (With Tools)42.2024 / 26Thinking (With Tools)+19.20
Terminal-Bench2725 / 35Normal (With Tools)2626 / 35Normal (With Tools)+1

General Knowledge

Claude Sonnet 4.5 3/3
BenchmarkClaude Sonnet 4.5Claude Sonnet 4Diff
MMLU Pro888 / 176Thinking (No Tools)8440 / 176Thinking (No Tools)+4
LiveBench53.6985 / 117Normal (No Tools)50.9891 / 117Normal (No Tools)+2.71
ARC-AGI-125.50127 / 147Normal (No Tools)23.80128 / 147Normal (No Tools)+1.70

Claw-style Agent Evaluation

Claude Sonnet 4.5 2/2
BenchmarkClaude Sonnet 4.5Claude Sonnet 4Diff
Claw Bench88.1013 / 29Thinking (With Tools)77.8023 / 29Thinking (With Tools)+10.30
Pinch Bench88.205 / 38Thinking (With Tools)80.5023 / 38Thinking (With Tools)+7.70

General Evaluation

Claude Sonnet 4.5 1/1
BenchmarkClaude Sonnet 4.5Claude Sonnet 4Diff
GPQA Diamond73.70292 / 462Normal (No Tools)68337 / 462Normal (No Tools)+5.70

Instruction Following

Claude Sonnet 4 1/1
BenchmarkClaude Sonnet 4.5Claude Sonnet 4Diff
IF Bench42.70199 / 282Normal (No Tools)45.40180 / 282Normal (No Tools)-2.70

Long Context

Claude Sonnet 4.5 1/1
BenchmarkClaude Sonnet 4.5Claude Sonnet 4Diff
AA-LCR54133 / 170Normal (No Tools)44150 / 170Normal (No Tools)+10

Multimodal Understanding

Claude Sonnet 4.5 1/1
BenchmarkClaude Sonnet 4.5Claude Sonnet 4Diff
MMMU-Pro65.20151 / 227Normal (No Tools)62.40163 / 227Normal (No Tools)+2.80

Productivity Knowledge

Claude Sonnet 4.5 1/1
BenchmarkClaude Sonnet 4.5Claude Sonnet 4Diff
GDPval-AA3910 / 15Thinking (No Tools)3313 / 15Thinking (No Tools)+6

Writing and Creative Capabilities

Claude Sonnet 4.5 1/1
BenchmarkClaude Sonnet 4.5Claude Sonnet 4Diff
Creative Writing1,67430 / 106Normal (No Tools)1,48053 / 106Normal (No Tools)+194.10

Specs

FieldClaude Sonnet 4.5Claude Sonnet 4
PublisherAnthropicAnthropic
Release date2025-09-302025-05-23
Model typeChat modelReasoning model
ArchitectureDenseDense
ParametersNot availableNot available
Context length1000K200K
Max output64K64K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemClaude Sonnet 4.5Claude Sonnet 4
Text input$3 / 1M tokens$3 / 1M tokens
Text output$15 / 1M tokens$15 / 1M tokens
Cache read$0.3 / 1M tokens$0.3 / 1M tokens
Cache write$3.75 / 1M tokens$3.75 / 1M tokens

Summary

  • Claude Sonnet 4.5leads in:Agent Level Benchmark (4/4), Coding and Software Engineer (4/4), AI Agent - Tool Usage (3/3), General Knowledge (3/3), Claw-style Agent Evaluation (2/2), General Evaluation (1/1), Long Context (1/1), Multimodal Understanding (1/1), Productivity Knowledge (1/1), Writing and Creative Capabilities (1/1)
  • Claude Sonnet 4leads in:Instruction Following (1/1)
  • Tied in:Math and Reasoning

On average across the 26 shared benchmarks, Claude Sonnet 4.5 scores 19.53 higher.

Largest single-benchmark gap: Creative Writing — Claude Sonnet 4.5 1,674 vs Claude Sonnet 4 1,480 (+194.10).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.