DataLearner 标志

Claude Sonnet 4.5vsClaude Sonnet 4

在 21 个共同 benchmark 中,Claude Sonnet 4.5 整体领先:Claude Sonnet 4.5 领先 18 项,Claude Sonnet 4 领先 1 项,持平 2 项,平均分差 +22.01。

Anthropic
Claude Sonnet 4.5

Anthropic · 2025-09-30 · 聊天大模型

Anthropic
Claude Sonnet 4

Anthropic · 2025-05-23 · 推理大模型

Claude Sonnet 4.518 (86%)持平2(5%)1 Claude Sonnet 4

评测分数

按能力类目分组,每组内按分差大小排列;共 21 项。

Coding and Software Engineer

Claude Sonnet 4.5 领先 4/4
评测项Claude Sonnet 4.5Claude Sonnet 4分差
CodeClash1,3891 / 8Normal (With Tools)1,2234 / 8Normal (With Tools)+166
LiveCodeBench5977 / 128Normal (No Tools)48.50170 / 250Normal (No Tools)+10.50
SWE-bench Verified828 / 116Parallel Thinking (With Tools)80.2014 / 116Parallel Thinking (With Tools)+1.80
SWE-Bench Pro - Public43.6054 / 62Thinking (No Tools)42.7055 / 62Thinking (No Tools)+0.90

General Knowledge

Claude Sonnet 4.5 领先 4/4
评测项Claude Sonnet 4.5Claude Sonnet 4分差
MMLU Pro887 / 134Thinking (No Tools)8439 / 175Thinking (No Tools)+4
LiveBench53.6985 / 117Normal (No Tools)50.9891 / 117Normal (No Tools)+2.71
ARC-AGI-125.50127 / 147Normal (No Tools)23.80128 / 147Normal (No Tools)+1.70
HLE7.10184 / 197Normal (No Tools)5.52467 / 563Normal (No Tools)+1.58

Math and Reasoning

胶着 4/4
评测项Claude Sonnet 4.5Claude Sonnet 4分差
FrontierMath5.2038 / 60Normal (No Tools)4.1041 / 60Normal (No Tools)+1.10
AIME202537107 / 117Normal (No Tools)38170 / 215Normal (No Tools)-1
IMO-ProofBench27.108 / 16Thinking (No Tools)27.108 / 16Thinking (No Tools)持平
IMO-ProofBench Advanced4.8019 / 24Thinking (No Tools)4.8019 / 24Thinking (No Tools)持平

AI Agent - Tool Usage

Claude Sonnet 4.5 领先 2/2
评测项Claude Sonnet 4.5Claude Sonnet 4分差
OSWorld-Verified61.4022 / 26Thinking (With Tools)42.2024 / 26Thinking (With Tools)+19.20
Terminal-Bench2725 / 35Normal (With Tools)2626 / 35Normal (With Tools)+1

Claw-style Agent Evaluation

Claude Sonnet 4.5 领先 2/2
评测项Claude Sonnet 4.5Claude Sonnet 4分差
Claw Bench88.1013 / 29Thinking (With Tools)77.8023 / 29Thinking (With Tools)+10.30
Pinch Bench88.205 / 38Thinking (With Tools)80.5023 / 38Thinking (With Tools)+7.70

Agent Level Benchmark

Claude Sonnet 4.5 领先 1/1
评测项Claude Sonnet 4.5Claude Sonnet 4分差
τ²-Bench7126 / 44Normal (With Tools)5235 / 44Normal (With Tools)+19

Long Context

Claude Sonnet 4.5 领先 1/1
评测项Claude Sonnet 4.5Claude Sonnet 4分差
AA-LCR54133 / 170Normal (No Tools)44150 / 170Normal (No Tools)+10

Productivity Knowledge

Claude Sonnet 4.5 领先 1/1
评测项Claude Sonnet 4.5Claude Sonnet 4分差
GDPval-AA3910 / 15Thinking (No Tools)3313 / 15Thinking (No Tools)+6

Writing and Creative Capabilities

Claude Sonnet 4.5 领先 1/1
评测项Claude Sonnet 4.5Claude Sonnet 4分差
Creative Writing1,67430 / 106Normal (No Tools)1,48053 / 106Normal (No Tools)+194.10

科学与综合推理

Claude Sonnet 4.5 领先 1/1
评测项Claude Sonnet 4.5Claude Sonnet 4分差
GPQA Diamond73.70179 / 274Normal (No Tools)68337 / 462Normal (No Tools)+5.70

规格对比

字段Claude Sonnet 4.5Claude Sonnet 4
发布机构AnthropicAnthropic
发布时间2025-09-302025-05-23
模型类型聊天大模型推理大模型
架构稠密模型稠密模型
参数规模暂无数据暂无数据
上下文长度1000K200K
最大输出64K64K

API 调用价格

价格优先使用 DataLearner 配置的 API 记录;缺失项不做推测。

价格项Claude Sonnet 4.5Claude Sonnet 4
文本输入$3 / 1M tokens$3 / 1M tokens
文本输出$15 / 1M tokens$15 / 1M tokens
缓存读取$0.3 / 1M tokens$0.3 / 1M tokens
缓存写入$3.75 / 1M tokens$3.75 / 1M tokens

小结

  • Claude Sonnet 4.5在以下类目领先:Coding and Software Engineer (4/4)、General Knowledge (4/4)、AI Agent - Tool Usage (2/2)、Claw-style Agent Evaluation (2/2)、Agent Level Benchmark (1/1)、Long Context (1/1)、Productivity Knowledge (1/1)、Writing and Creative Capabilities (1/1)、科学与综合推理 (1/1)
  • 胶着类目:Math and Reasoning

21 个共同 benchmark 上,Claude Sonnet 4.5 平均高出 22.01 分。

单项差距最大的 benchmark:Creative Writing — Claude Sonnet 4.5 1,674,Claude Sonnet 4 1,480(分差 +194.10)。

本页正文由结构化模型、价格与 benchmark 数据生成,不使用实时 LLM 撰写。