Claude Opus 4.8vsClaude Opus 4.6
Across 13 shared benchmarks, Claude Opus 4.8 leads overall: Claude Opus 4.8 wins 11, Claude Opus 4.6 wins 2, with 0 ties and an average score difference of +27.12.
Claude Opus 4.8
Anthropic · 2026-05-28 · Reasoning model
Claude Opus 4.6
Anthropic · 2026-02-05 · Reasoning model
Claude Opus 4.811 wins(85%)(15%)2 winsClaude Opus 4.6
Benchmark scores
Grouped by capability, sorted by largest gap within each. 13 shared benchmarks.
Coding and Software Engineer
Claude Opus 4.8 2/3| Benchmark | Claude Opus 4.8 | Claude Opus 4.6 | Diff |
|---|---|---|---|
| Text Arena (Coding) | 1,5459 / 35Normal (No Tools) | 1,5558 / 35Normal (No Tools) | -10.30 |
| SWE-bench Verified | 88.605 / 114Extended (with tools) | 80.8410 / 114Extended (with tools) | +7.76 |
| WeirdML v2 | 70.4518 / 52Normal (With Tools) | 65.9023 / 52Normal (With Tools) | +4.55 |
AI Agent - Tool Usage
Claude Opus 4.8 2/2| Benchmark | Claude Opus 4.8 | Claude Opus 4.6 | Diff |
|---|---|---|---|
| OSWorld-Verified | 83.404 / 26Extended (with tools) | 72.7016 / 26Extended (with tools) | +10.70 |
| MCP-Atlas | 82.206 / 38Deep Thinking (With Tools) | 76.8013 / 38Deep Thinking (With Tools) | +5.40 |
General Knowledge
Claude Opus 4.8 2/2| Benchmark | Claude Opus 4.8 | Claude Opus 4.6 | Diff |
|---|---|---|---|
| HLE | 57.908 / 181Extended (with tools) | 5320 / 181Extended (with tools, internet) | +4.90 |
| LiveBench | 78.794 / 115Deep Thinking (No Tools) | 76.338 / 115Thinking High (No Tools) | +2.46 |
Math and Reasoning
Claude Opus 4.8 2/2| Benchmark | Claude Opus 4.8 | Claude Opus 4.6 | Diff |
|---|---|---|---|
| FrontierMath Tier 4 v2 | 56.109 / 34最高(无工具) | 26.8318 / 34最高(无工具) | +29.27 |
| FrontierMath v2 | 809 / 34最高(无工具) | 65.9616 / 34最高(无工具) | +14.04 |
AI Agent - Information Search
Claude Opus 4.8 1/1| Benchmark | Claude Opus 4.8 | Claude Opus 4.6 | Diff |
|---|---|---|---|
| BrowseComp | 84.309 / 54Thinking High (With Tools + Internet) | 8411 / 54Thinking (With Tools + Internet) | +0.30 |
Commonsense Reasoning
Claude Opus 4.6 1/1| Benchmark | Claude Opus 4.8 | Claude Opus 4.6 | Diff |
|---|---|---|---|
| SimpleBench | 64.8010 / 67Normal (No Tools) | 67.609 / 67Normal (No Tools) | -2.80 |
General Evaluation
Claude Opus 4.8 1/1| Benchmark | Claude Opus 4.8 | Claude Opus 4.6 | Diff |
|---|---|---|---|
| GPQA Diamond | 93.6011 / 226Thinking High (No Tools) | 91.3128 / 226Extended (no tools) | +2.29 |
Productivity Knowledge
Claude Opus 4.8 1/1| Benchmark | Claude Opus 4.8 | Claude Opus 4.6 | Diff |
|---|---|---|---|
| GDPval-AA | 1,8901 / 21Extended (with tools) | 1,6063 / 21Extended (with tools, internet) | +284 |
Specs
| Field | Claude Opus 4.8 | Claude Opus 4.6 |
|---|---|---|
| Publisher | Anthropic | Anthropic |
| Release date | 2026-05-28 | 2026-02-05 |
| Model type | Reasoning model | Reasoning model |
| Architecture | Dense | Dense |
| Parameters | Not available | Not available |
| Context length | 1M | 1000K |
| Max output | 125K | 64K |
API pricing
Prices use DataLearner records when available; missing fields are not inferred.
| Item | Claude Opus 4.8 | Claude Opus 4.6 |
|---|---|---|
| Text input | $5 / 1M tokens | $0.5 / 1M tokens |
| Text output | $25 / 1M tokens | $25 / 1M tokens |
| Cache read | $0.5 / 1M tokens | $0.5 / 1M tokens |
| Cache write | $6.25 / 1M tokens | $10 / 1M tokens |
Summary
- Claude Opus 4.8leads in:Coding and Software Engineer (2/3), AI Agent - Tool Usage (2/2), General Knowledge (2/2), Math and Reasoning (2/2), AI Agent - Information Search (1/1), General Evaluation (1/1), Productivity Knowledge (1/1)
- Claude Opus 4.6leads in:Commonsense Reasoning (1/1)
On average across the 13 shared benchmarks, Claude Opus 4.8 scores 27.12 higher.
Largest single-benchmark gap: GDPval-AA — Claude Opus 4.8 1,890 vs Claude Opus 4.6 1,606 (+284).
Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.