DataLearner logo

GPT-5vsClaude Opus 4

Across 12 shared benchmarks, GPT-5 leads overall: GPT-5 wins 11, Claude Opus 4 wins 1, with 0 ties and an average score difference of +16.28.

OpenAI
GPT-5

OpenAI · 2025-08-07 · Foundation model

Anthropic
Claude Opus 4

Anthropic · 2025-05-23 · Reasoning model

GPT-511 wins(92%)(8%)1 winClaude Opus 4

Benchmark scores

Grouped by capability, sorted by largest gap within each. 12 shared benchmarks.

General Knowledge

GPT-5 4/4
BenchmarkGPT-5Claude Opus 4Diff
ARC-AGI65.7033 / 6835.7051 / 68+30
HLE35.2073 / 17210.70144 / 172+24.50
GPQA Diamond87.3040 / 18779.6085 / 187+7.70
ARC-AGI-29.9040 / 628.6042 / 62+1.30

Math and Reasoning

GPT-5 4/4
BenchmarkGPT-5Claude Opus 4Diff
IMO-ProofBench592 / 162.9016 / 16+56.10
AIME202599.609 / 10775.5066 / 107+24.10
FrontierMath24.8015 / 604.5039 / 60+20.30
FrontierMath - Tier 412.5029 / 80Thinking High (No Tools)4.2040 / 80+8.30

Agent Level Benchmark

GPT-5 2/2
BenchmarkGPT-5Claude Opus 4Diff
Aider-Polyglot881 / 59Thinking High (No Tools)70.7016 / 59Normal (No Tools)+17.30
τ²-Bench8015 / 4372.5023 / 43+7.50

Coding and Software Engineer

GPT-5 1/1
BenchmarkGPT-5Claude Opus 4Diff
SWE-bench Verified72.8050 / 11272.5052 / 112+0.30

Commonsense Reasoning

Claude Opus 4 1/1
BenchmarkGPT-5Claude Opus 4Diff
Simple Bench56.7020 / 63Thinking High (No Tools)58.8017 / 63Thinking (No Tools)-2.10

Specs

FieldGPT-5Claude Opus 4
PublisherOpenAIAnthropic
Release date2025-08-072025-05-23
Model typeFoundation modelReasoning model
ArchitectureDenseDense
ParametersNot availableNot available
Context length400K200K
Max output128K32K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGPT-5Claude Opus 4
Text input$1.25 / 1M tokens$15 / 1M tokens
Text output$10 / 1M tokens$75 / 1M tokens
Cache read$0.125 / 1M tokens$1.5 / 1M tokens
Cache write$0 / 1M tokens$18.75 / 1M tokens

Summary

  • GPT-5leads in:General Knowledge (4/4), Math and Reasoning (4/4), Agent Level Benchmark (2/2), Coding and Software Engineer (1/1)
  • Claude Opus 4leads in:Commonsense Reasoning (1/1)

On average across the 12 shared benchmarks, GPT-5 scores 16.28 higher.

Largest single-benchmark gap: IMO-ProofBench — GPT-5 59 vs Claude Opus 4 2.90 (+56.10).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.