DataLearner logo

Claude Opus 5vsGPT-5.6 Sol

Across 8 shared benchmarks, Claude Opus 5 leads overall: Claude Opus 5 wins 6, GPT-5.6 Sol wins 1, with 1 ties and an average score difference of +39.40.

Anthropic
Claude Opus 5

Anthropic · 2026-07-24 · Reasoning model

OpenAI
GPT-5.6 Sol

OpenAI · 2026-06-26 · Reasoning model

Claude Opus 56 wins(75%)Ties1(13%)1 winGPT-5.6 Sol

Benchmark scores

Grouped by capability, sorted by largest gap within each. 8 shared benchmarks.

AI Agent - Tool Usage

Claude Opus 5 2/2
BenchmarkClaude Opus 5GPT-5.6 SolDiff
Terminal-Bench 4.051.823 / 13Max (With Tools)37.276 / 13Max (With Tools)+14.55
Terminal-Bench-Science 0.1303 / 11Max (With Tools)22.404 / 11Max (With Tools)+7.60

Productivity Knowledge

Claude Opus 5 2/2
BenchmarkClaude Opus 5GPT-5.6 SolDiff
GDPval-AA v21,8611 / 26Max (With Tools)1,7288 / 26Max (With Tools)+133
AA-AnalystAgent53.752 / 12Max (With Tools + Internet)47.505 / 12Max (With Tools + Internet)+6.25

General Knowledge

Even 1/1
BenchmarkClaude Opus 5GPT-5.6 SolDiff
ARC-AGI-197.504 / 92Extra-High (No Tools)97.504 / 92Extra-High (No Tools)

Math and Reasoning

GPT-5.6 Sol 1/1
BenchmarkClaude Opus 5GPT-5.6 SolDiff
FrontierMath v285.615 / 58Max (No Tools)89.121 / 58Max (No Tools)-3.51

Text Embedding

Claude Opus 5 1/1
BenchmarkClaude Opus 5GPT-5.6 SolDiff
Context Arena97.721 / 126Max (No Tools)97.632 / 126Max (No Tools)+0.09

Writing and Creative Capabilities

Claude Opus 5 1/1
BenchmarkClaude Opus 5GPT-5.6 SolDiff
Creative Writing2,1213 / 106Normal (No Tools)1,9636 / 106Normal (No Tools)+157.20

Specs

FieldClaude Opus 5GPT-5.6 Sol
PublisherAnthropicOpenAI
Release date2026-07-242026-06-26
Model typeReasoning modelReasoning model
ArchitectureDenseDense
ParametersNot availableNot available
Context length1M1.05M
Max output128K128K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemClaude Opus 5GPT-5.6 Sol
Text input$5 / 1M tokens$4 / 1M tokens
Text output$25 / 1M tokens$20 / 1M tokens
Cache read$0.5 / 1M tokens$0.4 / 1M tokens
Cache write$6.25 / 1M tokens$5 / 1M tokens

Summary

  • Claude Opus 5leads in:AI Agent - Tool Usage (2/2), Productivity Knowledge (2/2), Text Embedding (1/1), Writing and Creative Capabilities (1/1)
  • GPT-5.6 Solleads in:Math and Reasoning (1/1)
  • Tied in:General Knowledge

On average across the 8 shared benchmarks, Claude Opus 5 scores 39.40 higher.

Largest single-benchmark gap: Creative Writing — Claude Opus 5 2,121 vs GPT-5.6 Sol 1,963 (+157.20).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.