DataLearner logo

GPT-5.1vsClaude Opus 4

Across 5 shared benchmarks, Claude Opus 4 leads overall: GPT-5.1 wins 1, Claude Opus 4 wins 4, with 0 ties and an average score difference of -11.18.

OpenAI
GPT-5.1

OpenAI · 2025-11-12 · Reasoning model

Anthropic
Claude Opus 4

Anthropic · 2025-05-23 · Reasoning model

GPT-5.11 win(20%)(80%)4 winsClaude Opus 4

Benchmark scores

Grouped by capability, sorted by largest gap within each. 5 shared benchmarks.

Math and Reasoning

Even 2/2
BenchmarkGPT-5.1Claude Opus 4Diff
AIME202538170 / 215Normal (No Tools)75.50106 / 215Normal (No Tools)-37.50
IMO-ProofBench Advanced7.1016 / 24Thinking (No Tools)2.9024 / 24Thinking (No Tools)+4.20

Coding and Software Engineer

Claude Opus 4 1/1
BenchmarkGPT-5.1Claude Opus 4Diff
LiveCodeBench49.40167 / 250Normal (No Tools)56.60143 / 250Normal (No Tools)-7.20

General Evaluation

Claude Opus 4 1/1
BenchmarkGPT-5.1Claude Opus 4Diff
GPQA Diamond64.30361 / 462Normal (No Tools)79.60241 / 462Normal (No Tools)-15.30

Instruction Following

Claude Opus 4 1/1
BenchmarkGPT-5.1Claude Opus 4Diff
IF Bench43.20192 / 282Normal (No Tools)43.30191 / 282Normal (No Tools)-0.10

Specs

FieldGPT-5.1Claude Opus 4
PublisherOpenAIAnthropic
Release date2025-11-122025-05-23
Model typeReasoning modelReasoning model
ArchitectureDenseDense
ParametersNot availableNot available
Context length400K200K
Max output128K32K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGPT-5.1Claude Opus 4
Text input$1.25 / 1M tokens$15 / 1M tokens
Text output$10 / 1M tokens$75 / 1M tokens
Cache read$0.125 / 1M tokens$1.5 / 1M tokens
Cache write$0 / 1M tokens$18.75 / 1M tokens

Summary

  • Claude Opus 4leads in:Coding and Software Engineer (1/1), General Evaluation (1/1), Instruction Following (1/1)
  • Tied in:Math and Reasoning

On average across the 5 shared benchmarks, Claude Opus 4 scores 11.18 higher.

Largest single-benchmark gap: AIME2025 — GPT-5.1 38 vs Claude Opus 4 75.50 (-37.50).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.