DataLearner logo

Opus 4.7vsClaude Opus 4.6

Across 22 shared benchmarks, Opus 4.7 leads overall: Opus 4.7 wins 18, Claude Opus 4.6 wins 4, with 0 ties and an average score difference of +7.34.

Anthropic
Opus 4.7

Anthropic · 2026-04-16 · Reasoning model

Anthropic
Claude Opus 4.6

Anthropic · 2026-02-05 · Reasoning model

Opus 4.718 wins(82%)(18%)4 winsClaude Opus 4.6

Benchmark scores

Grouped by capability, sorted by largest gap within each. 22 shared benchmarks.

Coding and Software Engineer

Opus 4.7 4/4
BenchmarkOpus 4.7Claude Opus 4.6Diff
GSO44.121 / 21Normal (With Tools)33.335 / 21Normal (With Tools)+10.79
WeirdML v276.4013 / 52Normal (With Tools)65.9023 / 52Normal (With Tools)+10.50
Text Arena (Coding)1,5627 / 35Normal (No Tools)1,5558 / 35Normal (No Tools)+7.04
SWE-bench Verified87.606 / 116Extended (with tools)80.8410 / 116Extended (with tools)+6.76

Agent Level Benchmark

Opus 4.7 2/3
BenchmarkOpus 4.7Claude Opus 4.6Diff
τ³-Banking40.2130 / 164Max (With Tools)27.3269 / 164Max (With Tools)+12.89
τ²-Bench - Telecom74128 / 264Normal (With Tools)84.8090 / 264Normal (With Tools)-10.80
Terminal Bench Hard54.5015 / 244Normal (With Tools)48.5025 / 244Normal (With Tools)+6

AI Agent - Tool Usage

Opus 4.7 3/3
BenchmarkOpus 4.7Claude Opus 4.6Diff
OSWorld-Verified7811 / 26Extended (with tools)72.7016 / 26Extended (with tools)+5.30
Terminal Bench 2.069.406 / 48Extended (with tools)65.4011 / 48Extended (with tools)+4
MCP-Atlas79.1013 / 43Deep Thinking (With Tools)76.8018 / 43Deep Thinking (With Tools)+2.30

General Knowledge

Opus 4.7 2/2
BenchmarkOpus 4.7Claude Opus 4.6Diff
HLE33.30184 / 563Normal (No Tools) · Text only19.10299 / 563Normal (No Tools) · Text only+14.20
CritPt5.1094 / 200Normal (No Tools)2.80120 / 200Normal (No Tools)+2.30

Math and Reasoning

Opus 4.7 2/2
BenchmarkOpus 4.7Claude Opus 4.6Diff
FrontierMath Tier 4 v231.7117 / 42Max (No Tools)26.8323 / 42Max (No Tools)+4.88
FrontierMath v270.1815 / 58Max (No Tools)65.9618 / 58Max (No Tools)+4.21

Claw-style Agent Evaluation

Opus 4.7 1/1
BenchmarkOpus 4.7Claude Opus 4.6Diff
PinchBench v276.0115 / 45Reported best (effort unspecified)69.9026 / 45Reported best (effort unspecified)+6.11

Commonsense Reasoning

Claude Opus 4.6 1/1
BenchmarkOpus 4.7Claude Opus 4.6Diff
SimpleBench61.7027 / 93Normal (No Tools)67.6019 / 93Normal (No Tools)-5.90

General Evaluation

Opus 4.7 1/1
BenchmarkOpus 4.7Claude Opus 4.6Diff
GPQA Diamond88.50110 / 462Normal (No Tools)84176 / 462Normal (No Tools)+4.50

Instruction Following

Claude Opus 4.6 1/1
BenchmarkOpus 4.7Claude Opus 4.6Diff
IF Bench43.60189 / 282Normal (No Tools)44.60182 / 282Normal (No Tools)-1

Long Context

Opus 4.7 1/1
BenchmarkOpus 4.7Claude Opus 4.6Diff
AA-LCR75.7082 / 170Normal (No Tools)67116 / 170Normal (No Tools)+8.70

Multimodal Understanding

Opus 4.7 1/1
BenchmarkOpus 4.7Claude Opus 4.6Diff
MMMU-Pro76.4073 / 227Normal (No Tools)72.50114 / 227Normal (No Tools)+3.90

Text Embedding

Claude Opus 4.6 1/1
BenchmarkOpus 4.7Claude Opus 4.6Diff
Context Arena22.12122 / 126Normal (No Tools)60.3376 / 126Normal (No Tools)-38.21

Writing and Creative Capabilities

Opus 4.7 1/1
BenchmarkOpus 4.7Claude Opus 4.6Diff
Creative Writing1,9079 / 106Normal (No Tools)1,80419 / 106Normal (No Tools)+103

Specs

FieldOpus 4.7Claude Opus 4.6
PublisherAnthropicAnthropic
Release date2026-04-162026-02-05
Model typeReasoning modelReasoning model
ArchitectureDenseDense
ParametersNot availableNot available
Context length1000K1000K
Max output128K64K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemOpus 4.7Claude Opus 4.6
Text input$5 / 1M tokens$5 / 1M tokens
Text output$25 / 1M tokens$25 / 1M tokens
Cache read$0.5 / 1M tokens$0.5 / 1M tokens
Cache write$6.25 / 1M tokens$6.25 / 1M tokens

Summary

  • Opus 4.7leads in:Coding and Software Engineer (4/4), Agent Level Benchmark (2/3), AI Agent - Tool Usage (3/3), General Knowledge (2/2), Math and Reasoning (2/2), Claw-style Agent Evaluation (1/1), General Evaluation (1/1), Long Context (1/1), Multimodal Understanding (1/1), Writing and Creative Capabilities (1/1)
  • Claude Opus 4.6leads in:Commonsense Reasoning (1/1), Instruction Following (1/1), Text Embedding (1/1)

On average across the 22 shared benchmarks, Opus 4.7 scores 7.34 higher.

Largest single-benchmark gap: Creative Writing — Opus 4.7 1,907 vs Claude Opus 4.6 1,804 (+103).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.