DataLearner logo

Claude Sonnet 5vsClaude Sonnet 4.5

Across 10 shared benchmarks, Claude Sonnet 5 leads overall: Claude Sonnet 5 wins 10, Claude Sonnet 4.5 wins 0, with 0 ties and an average score difference of +36.60.

Anthropic
Claude Sonnet 5

Anthropic · 2026-06-30 · Multimodal model

Anthropic
Claude Sonnet 4.5

Anthropic · 2025-09-30 · Chat model

Claude Sonnet 510 wins(100%)(0%)0 winsClaude Sonnet 4.5

Benchmark scores

Grouped by capability, sorted by largest gap within each. 10 shared benchmarks.

Coding and Software Engineer

Claude Sonnet 5 3/3
BenchmarkClaude Sonnet 5Claude Sonnet 4.5Diff
Text Arena (Coding)1,54410 / 35Thinking High (No Tools)1,38732 / 35Normal (No Tools)+157.16
WeirdML v268.7820 / 52Thinking High (With Tools)46.7040 / 52Normal (With Tools)+22.08
SWE-bench Verified85.207 / 114极高强度思考(工具)828 / 114+3.20

Math and Reasoning

Claude Sonnet 5 2/2
BenchmarkClaude Sonnet 5Claude Sonnet 4.5Diff
FrontierMath v265.6117 / 34最高(无工具)23.8631 / 34Thinking (No Tools, 32K Budget)+41.75
FrontierMath Tier 4 v229.2716 / 34最高(无工具)2.4430 / 34Thinking (No Tools, 32K Budget)+26.83

AI Agent - Information Search

Claude Sonnet 5 1/1
BenchmarkClaude Sonnet 5Claude Sonnet 4.5Diff
BrowseComp84.707 / 54Thinking (With Tools + Internet)24.1052 / 54+60.60

AI Agent - Tool Usage

Claude Sonnet 5 1/1
BenchmarkClaude Sonnet 5Claude Sonnet 4.5Diff
OSWorld-Verified81.206 / 26极高强度思考(工具)61.4022 / 26+19.80

Commonsense Reasoning

Claude Sonnet 5 1/1
BenchmarkClaude Sonnet 5Claude Sonnet 4.5Diff
SimpleBench57.9023 / 67Normal (No Tools)54.3027 / 67Normal (No Tools)+3.60

General Evaluation

Claude Sonnet 5 1/1
BenchmarkClaude Sonnet 5Claude Sonnet 4.5Diff
GPQA Diamond90.5333 / 226极高强度思考(无工具)83.4099 / 226+7.13

General Knowledge

Claude Sonnet 5 1/1
BenchmarkClaude Sonnet 5Claude Sonnet 4.5Diff
HLE57.409 / 181极高强度思考(工具)33.6085 / 181+23.80

Specs

FieldClaude Sonnet 5Claude Sonnet 4.5
PublisherAnthropicAnthropic
Release date2026-06-302025-09-30
Model typeMultimodal modelChat model
ArchitectureDenseDense
ParametersNot availableNot available
Context length1M1000K
Max output128K64K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemClaude Sonnet 5Claude Sonnet 4.5
Text input$2 / 1M tokens$3 / 1M tokens
Text output$10 / 1M tokens$15 / 1M tokens
Cache read$0.3 / 1M tokens$0.3 / 1M tokens
Cache write$3.75 / 1M tokens$3.75 / 1M tokens

Summary

  • Claude Sonnet 5leads in:Coding and Software Engineer (3/3), Math and Reasoning (2/2), AI Agent - Information Search (1/1), AI Agent - Tool Usage (1/1), Commonsense Reasoning (1/1), General Evaluation (1/1), General Knowledge (1/1)

On average across the 10 shared benchmarks, Claude Sonnet 5 scores 36.60 higher.

Largest single-benchmark gap: Text Arena (Coding) — Claude Sonnet 5 1,544 vs Claude Sonnet 4.5 1,387 (+157.16).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.