DataLearner logo

Claude Opus 5vsClaude Opus 4.8

Across 5 shared benchmarks, Claude Opus 5 leads overall: Claude Opus 5 wins 5, Claude Opus 4.8 wins 0, with 0 ties and an average score difference of +8.10.

Anthropic
Claude Opus 5

Anthropic · 2026-07-24 · Reasoning model

Anthropic
Claude Opus 4.8

Anthropic · 2026-05-28 · Reasoning model

Claude Opus 55 wins(100%)(0%)0 winsClaude Opus 4.8

Benchmark scores

Grouped by capability, sorted by largest gap within each. 5 shared benchmarks.

Coding and Software Engineer

Claude Opus 5 3/3
BenchmarkClaude Opus 5Claude Opus 4.8Diff
SWE-Bench Pro - Public79.202 / 54Max (With Tools)69.204 / 54Extended (with tools)+10
DeepSWE68.804 / 19Max (With Tools)598 / 19Deep Thinking (With Tools)+9.80
SWE-bench Verified961 / 112Max (With Tools)88.605 / 112Extended (with tools)+7.40

AI Agent - Information Search

Claude Opus 5 1/1
BenchmarkClaude Opus 5Claude Opus 4.8Diff
BrowseComp90.802 / 53Max (With Tools + Internet)84.309 / 53Thinking High (With Tools + Internet)+6.50

General Knowledge

Claude Opus 5 1/1
BenchmarkClaude Opus 5Claude Opus 4.8Diff
HLE64.701 / 172Max (With Tools)57.907 / 172Extended (with tools)+6.80

Specs

FieldClaude Opus 5Claude Opus 4.8
PublisherAnthropicAnthropic
Release date2026-07-242026-05-28
Model typeReasoning modelReasoning model
ArchitectureDenseDense
ParametersNot availableNot available
Context length1M1M
Max output128K125K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemClaude Opus 5Claude Opus 4.8
Text input$5 / 1M tokens$5 / 1M tokens
Text output$25 / 1M tokens$25 / 1M tokens
Cache read$0.5 / 1M tokens$0.5 / 1M tokens
Cache write$6.25 / 1M tokens$6.25 / 1M tokens

Summary

  • Claude Opus 5leads in:Coding and Software Engineer (3/3), AI Agent - Information Search (1/1), General Knowledge (1/1)

On average across the 5 shared benchmarks, Claude Opus 5 scores 8.10 higher.

Largest single-benchmark gap: SWE-Bench Pro - Public — Claude Opus 5 79.20 vs Claude Opus 4.8 69.20 (+10).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.