DataLearner logo

Claude Opus 5vsClaude Opus 4.6

Across 7 shared benchmarks, Claude Opus 5 leads overall: Claude Opus 5 wins 7, Claude Opus 4.6 wins 0, with 0 ties and an average score difference of +15.85.

Anthropic
Claude Opus 5

Anthropic · 2026-07-24 · Reasoning model

Anthropic
Claude Opus 4.6

Anthropic · 2026-02-05 · Reasoning model

Claude Opus 57 wins(100%)(0%)0 winsClaude Opus 4.6

Benchmark scores

Grouped by capability, sorted by largest gap within each. 7 shared benchmarks.

General Knowledge

Claude Opus 5 4/4
BenchmarkClaude Opus 5Claude Opus 4.6Diff
ARC-AGI-330.201 / 9Thinking High (No Tools)04 / 9最高(无工具)+30.20
ARC-AGI-290.401 / 62最高(无工具)66.3017 / 62Extended (no tools)+24.10
HLE64.701 / 172Max (With Tools)5318 / 172Extended (with tools, internet)+11.70
ARC-AGI97.501 / 68极高强度思考(无工具)9213 / 68Extended (no tools)+5.50

Coding and Software Engineer

Claude Opus 5 2/2
BenchmarkClaude Opus 5Claude Opus 4.6Diff
SWE-bench Multilingual89.501 / 23Max (With Tools)7214 / 23Extended (with tools)+17.50
SWE-bench Verified961 / 112Max (With Tools)80.8410 / 112Extended (with tools)+15.16

AI Agent - Information Search

Claude Opus 5 1/1
BenchmarkClaude Opus 5Claude Opus 4.6Diff
BrowseComp90.802 / 53Max (With Tools + Internet)8411 / 53Thinking (With Tools + Internet)+6.80

Specs

FieldClaude Opus 5Claude Opus 4.6
PublisherAnthropicAnthropic
Release date2026-07-242026-02-05
Model typeReasoning modelReasoning model
ArchitectureDenseDense
ParametersNot availableNot available
Context length1M1000K
Max output128K64K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemClaude Opus 5Claude Opus 4.6
Text input$5 / 1M tokens$0.5 / 1M tokens
Text output$25 / 1M tokens$25 / 1M tokens
Cache read$0.5 / 1M tokens$0.5 / 1M tokens
Cache write$6.25 / 1M tokens$10 / 1M tokens

Summary

  • Claude Opus 5leads in:General Knowledge (4/4), Coding and Software Engineer (2/2), AI Agent - Information Search (1/1)

On average across the 7 shared benchmarks, Claude Opus 5 scores 15.85 higher.

Largest single-benchmark gap: ARC-AGI-3 — Claude Opus 5 30.20 vs Claude Opus 4.6 0 (+30.20).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.