DataLearner logo

GPT-5vsClaude Opus 4

Across 9 shared benchmarks, GPT-5 leads overall: GPT-5 wins 5, Claude Opus 4 wins 4, with 0 ties and an average score difference of +9.31.

OpenAI
GPT-5

OpenAI · 2025-08-07 · Foundation model

Anthropic
Claude Opus 4

Anthropic · 2025-05-23 · Reasoning model

GPT-55 wins(56%)(44%)4 winsClaude Opus 4

Benchmark scores

Grouped by capability, sorted by largest gap within each. 9 shared benchmarks.

Math and Reasoning

GPT-5 2/3
BenchmarkGPT-5Claude Opus 4Diff
IMO-ProofBench592 / 16Thinking (No Tools)2.9016 / 16Thinking (No Tools)+56.10
IMO-ProofBench Advanced2011 / 24Thinking (No Tools)2.9024 / 24Thinking (No Tools)+17.10
AIME202561.90135 / 215Normal (No Tools)75.50106 / 215Normal (No Tools)-13.60

Agent Level Benchmark

GPT-5 1/1
BenchmarkGPT-5Claude Opus 4Diff
τ²-Bench8015 / 44Thinking (With Tools)72.5024 / 44Thinking (With Tools)+7.50

Coding and Software Engineer

Claude Opus 4 1/1
BenchmarkGPT-5Claude Opus 4Diff
LiveCodeBench55.80149 / 250Normal (No Tools)56.60143 / 250Normal (No Tools)-0.80

General Evaluation

Claude Opus 4 1/1
BenchmarkGPT-5Claude Opus 4Diff
GPQA Diamond77.80257 / 462Normal (No Tools)79.60241 / 462Normal (No Tools)-1.80

General Knowledge

Claude Opus 4 1/1
BenchmarkGPT-5Claude Opus 4Diff
ARC-AGI-16142 / 147Normal (No Tools)35.67119 / 147Normal (No Tools)-29.67

Instruction Following

GPT-5 1/1
BenchmarkGPT-5Claude Opus 4Diff
IF Bench45.60178 / 282Normal (No Tools)43.30191 / 282Normal (No Tools)+2.30

Writing and Creative Capabilities

GPT-5 1/1
BenchmarkGPT-5Claude Opus 4Diff
Creative Writing1,62436 / 106Normal (No Tools)1,57741 / 106Normal (No Tools)+46.70

Specs

FieldGPT-5Claude Opus 4
PublisherOpenAIAnthropic
Release date2025-08-072025-05-23
Model typeFoundation modelReasoning model
ArchitectureDenseDense
ParametersNot availableNot available
Context length400K200K
Max output128K32K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGPT-5Claude Opus 4
Text input$1.25 / 1M tokens$15 / 1M tokens
Text output$10 / 1M tokens$75 / 1M tokens
Cache read$0.125 / 1M tokens$1.5 / 1M tokens
Cache write$0 / 1M tokens$18.75 / 1M tokens

Summary

  • GPT-5leads in:Math and Reasoning (2/3), Agent Level Benchmark (1/1), Instruction Following (1/1), Writing and Creative Capabilities (1/1)
  • Claude Opus 4leads in:Coding and Software Engineer (1/1), General Evaluation (1/1), General Knowledge (1/1)

On average across the 9 shared benchmarks, GPT-5 scores 9.31 higher.

Largest single-benchmark gap: IMO-ProofBench — GPT-5 59 vs Claude Opus 4 2.90 (+56.10).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.