DataLearner logo

Claude Opus 4.8vsClaude Opus 4.6

Across 13 shared benchmarks, Claude Opus 4.8 leads overall: Claude Opus 4.8 wins 11, Claude Opus 4.6 wins 2, with 0 ties and an average score difference of +27.12.

Anthropic
Claude Opus 4.8

Anthropic · 2026-05-28 · Reasoning model

Anthropic
Claude Opus 4.6

Anthropic · 2026-02-05 · Reasoning model

Claude Opus 4.811 wins(85%)(15%)2 winsClaude Opus 4.6

Benchmark scores

Grouped by capability, sorted by largest gap within each. 13 shared benchmarks.

Coding and Software Engineer

Claude Opus 4.8 2/3
BenchmarkClaude Opus 4.8Claude Opus 4.6Diff
Text Arena (Coding)1,5459 / 35Normal (No Tools)1,5558 / 35Normal (No Tools)-10.30
SWE-bench Verified88.605 / 114Extended (with tools)80.8410 / 114Extended (with tools)+7.76
WeirdML v270.4518 / 52Normal (With Tools)65.9023 / 52Normal (With Tools)+4.55

AI Agent - Tool Usage

Claude Opus 4.8 2/2
BenchmarkClaude Opus 4.8Claude Opus 4.6Diff
OSWorld-Verified83.404 / 26Extended (with tools)72.7016 / 26Extended (with tools)+10.70
MCP-Atlas82.206 / 38Deep Thinking (With Tools)76.8013 / 38Deep Thinking (With Tools)+5.40

General Knowledge

Claude Opus 4.8 2/2
BenchmarkClaude Opus 4.8Claude Opus 4.6Diff
HLE57.908 / 181Extended (with tools)5320 / 181Extended (with tools, internet)+4.90
LiveBench78.794 / 115Deep Thinking (No Tools)76.338 / 115Thinking High (No Tools)+2.46

Math and Reasoning

Claude Opus 4.8 2/2
BenchmarkClaude Opus 4.8Claude Opus 4.6Diff
FrontierMath Tier 4 v256.109 / 34最高(无工具)26.8318 / 34最高(无工具)+29.27
FrontierMath v2809 / 34最高(无工具)65.9616 / 34最高(无工具)+14.04

AI Agent - Information Search

Claude Opus 4.8 1/1
BenchmarkClaude Opus 4.8Claude Opus 4.6Diff
BrowseComp84.309 / 54Thinking High (With Tools + Internet)8411 / 54Thinking (With Tools + Internet)+0.30

Commonsense Reasoning

Claude Opus 4.6 1/1
BenchmarkClaude Opus 4.8Claude Opus 4.6Diff
SimpleBench64.8010 / 67Normal (No Tools)67.609 / 67Normal (No Tools)-2.80

General Evaluation

Claude Opus 4.8 1/1
BenchmarkClaude Opus 4.8Claude Opus 4.6Diff
GPQA Diamond93.6011 / 226Thinking High (No Tools)91.3128 / 226Extended (no tools)+2.29

Productivity Knowledge

Claude Opus 4.8 1/1
BenchmarkClaude Opus 4.8Claude Opus 4.6Diff
GDPval-AA1,8901 / 21Extended (with tools)1,6063 / 21Extended (with tools, internet)+284

Specs

FieldClaude Opus 4.8Claude Opus 4.6
PublisherAnthropicAnthropic
Release date2026-05-282026-02-05
Model typeReasoning modelReasoning model
ArchitectureDenseDense
ParametersNot availableNot available
Context length1M1000K
Max output125K64K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemClaude Opus 4.8Claude Opus 4.6
Text input$5 / 1M tokens$0.5 / 1M tokens
Text output$25 / 1M tokens$25 / 1M tokens
Cache read$0.5 / 1M tokens$0.5 / 1M tokens
Cache write$6.25 / 1M tokens$10 / 1M tokens

Summary

  • Claude Opus 4.8leads in:Coding and Software Engineer (2/3), AI Agent - Tool Usage (2/2), General Knowledge (2/2), Math and Reasoning (2/2), AI Agent - Information Search (1/1), General Evaluation (1/1), Productivity Knowledge (1/1)
  • Claude Opus 4.6leads in:Commonsense Reasoning (1/1)

On average across the 13 shared benchmarks, Claude Opus 4.8 scores 27.12 higher.

Largest single-benchmark gap: GDPval-AA — Claude Opus 4.8 1,890 vs Claude Opus 4.6 1,606 (+284).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.