DataLearner logo

Claude Opus 4.8vsOpus 4.7

Across 14 shared benchmarks, Claude Opus 4.8 leads overall: Claude Opus 4.8 wins 11, Opus 4.7 wins 3, with 0 ties and an average score difference of +3.28.

Anthropic
Claude Opus 4.8

Anthropic · 2026-05-28 · Reasoning model

Anthropic
Opus 4.7

Anthropic · 2026-04-16 · Reasoning model

Claude Opus 4.811 wins(79%)(21%)3 winsOpus 4.7

Benchmark scores

Grouped by capability, sorted by largest gap within each. 14 shared benchmarks.

Coding and Software Engineer

Even 4/4
BenchmarkClaude Opus 4.8Opus 4.7Diff
Text Arena (Coding)1,5459 / 35Normal (No Tools)1,5627 / 35Normal (No Tools)-17.34
WeirdML v270.4518 / 52Normal (With Tools)76.4013 / 52Normal (With Tools)-5.95
SWE-Bench Pro - Public69.204 / 57Extended (with tools)64.308 / 57Extended (with tools)+4.90
SWE-bench Verified88.605 / 114Extended (with tools)87.606 / 114Extended (with tools)+1

AI Agent - Tool Usage

Claude Opus 4.8 3/3
BenchmarkClaude Opus 4.8Opus 4.7Diff
Terminal-Bench 2.178.9020 / 44Thinking High (With Tools)69.7032 / 44Thinking High (With Tools)+9.20
OSWorld-Verified83.404 / 26Extended (with tools)7811 / 26Extended (with tools)+5.40
MCP-Atlas82.206 / 38Deep Thinking (With Tools)79.109 / 38Deep Thinking (With Tools)+3.10

General Knowledge

Claude Opus 4.8 2/2
BenchmarkClaude Opus 4.8Opus 4.7Diff
HLE57.908 / 181Extended (with tools)54.7015 / 181Extended (with tools)+3.20
LiveBench78.794 / 115Deep Thinking (No Tools)76.917 / 115Deep Thinking (No Tools)+1.88

Math and Reasoning

Claude Opus 4.8 2/2
BenchmarkClaude Opus 4.8Opus 4.7Diff
FrontierMath Tier 4 v256.109 / 34最高(无工具)31.7114 / 34最高(无工具)+24.39
FrontierMath v2809 / 34最高(无工具)70.1814 / 34最高(无工具)+9.82

AI Agent - Information Search

Claude Opus 4.8 1/1
BenchmarkClaude Opus 4.8Opus 4.7Diff
BrowseComp84.309 / 54Thinking High (With Tools + Internet)79.3017 / 54Extended (with tools)+5

Commonsense Reasoning

Claude Opus 4.8 1/1
BenchmarkClaude Opus 4.8Opus 4.7Diff
SimpleBench64.8010 / 67Normal (No Tools)62.9012 / 67Normal (No Tools)+1.90

General Evaluation

Opus 4.7 1/1
BenchmarkClaude Opus 4.8Opus 4.7Diff
GPQA Diamond93.6011 / 226Thinking High (No Tools)94.205 / 226Extended (no tools)-0.60

Specs

FieldClaude Opus 4.8Opus 4.7
PublisherAnthropicAnthropic
Release date2026-05-282026-04-16
Model typeReasoning modelReasoning model
ArchitectureDenseDense
ParametersNot availableNot available
Context length1M1000K
Max output125K128K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemClaude Opus 4.8Opus 4.7
Text input$5 / 1M tokens$5 / 1M tokens
Text output$25 / 1M tokens$25 / 1M tokens
Cache read$0.5 / 1M tokens$0.5 / 1M tokens
Cache write$6.25 / 1M tokens$6.25 / 1M tokens

Summary

  • Claude Opus 4.8leads in:AI Agent - Tool Usage (3/3), General Knowledge (2/2), Math and Reasoning (2/2), AI Agent - Information Search (1/1), Commonsense Reasoning (1/1)
  • Opus 4.7leads in:General Evaluation (1/1)
  • Tied in:Coding and Software Engineer

On average across the 14 shared benchmarks, Claude Opus 4.8 scores 3.28 higher.

Largest single-benchmark gap: FrontierMath Tier 4 v2 — Claude Opus 4.8 56.10 vs Opus 4.7 31.71 (+24.39).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.