See key specs and per-benchmark scores for each model/mode. Scroll horizontally for all columns. 当前对比 2 个模型的评测数据与核心参数。

Opus 4.5
Anthropic
Best overall
Claude Opus 5 · 87.15
Best single
Claude Opus 5 · ARC-AGI 97.50
Modality coverage
Opus 4.5 · 2 modalities
Head to head
4
Benchmarks
0
Wins
4
Losses
-26.73
Average diff
Compare benchmark results across thinking modes and tool usage.
Data sourced primarily from official releases (GitHub, Hugging Face, papers), then benchmark leaderboards, then third-party evaluators. Learn about our data methodology
Complete scores for each model/mode across selected benchmarks.
4 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.
| Benchmark | Opus 4.5 | Claude Opus 5 |
|---|---|---|
ARC-AGI 综合评估 | 80.00Extended Thinking | 97.50Thinking Level · Extra High |
ARC-AGI-2 综合评估 | 37.60Extended Thinking | 90.40Thinking Level · High |
HLE 综合评估 | 43.20Extended Thinking | Tools | 64.70Thinking Level · High | Tools |
SWE-bench Verified 编程与软件工程 | 80.90Extended Thinking | Tools | 96.00Thinking Level · High | Tools |
Side-by-side input/output token pricing
Licensing, MoE architecture, and multi-modality support.
| Features & specs | Opus 4.5Anthropic | Claude Opus 5Anthropic |
|---|---|---|
Core specsRelease | 2025-11-25 | 2026-07-24 |
Context length | 200K | 1M |
Max output | 65536 | 128000 |
MoE | No | No |
LicenseCode Open Source | Not provided | Not provided |
Weights Open Source | Not provided | Not provided |
Commercial use | 不开源 | 不开源 |
Modality supportText Input/Output | / | / |
Image Input/Output | / | / |
ResourcesPaper / report | Introducing Claude Opus 4.5 | Introducing Claude Opus 5 |

Claude Opus 5
Anthropic