DataLearner logo

GPT-5.4vsClaude Opus 4.6

Across 13 shared benchmarks, Claude Opus 4.6 leads overall: GPT-5.4 wins 4, Claude Opus 4.6 wins 9, with 0 ties and an average score difference of -4.05.

OpenAI
GPT-5.4

OpenAI · 2026-03-05 · Multimodal model

Anthropic
Claude Opus 4.6

Anthropic · 2026-02-05 · Reasoning model

GPT-5.44 wins(31%)(69%)9 winsClaude Opus 4.6

Benchmark scores

Grouped by capability, sorted by largest gap within each. 13 shared benchmarks.

Agent Level Benchmark

Claude Opus 4.6 2/2
BenchmarkGPT-5.4Claude Opus 4.6Diff
τ²-Bench - Telecom64.30151 / 264Normal (With Tools)84.8090 / 264Normal (With Tools)-20.50
Terminal Bench Hard37.9059 / 244Normal (With Tools)48.5025 / 244Normal (With Tools)-10.60

Claw-style Agent Evaluation

GPT-5.4 2/2
BenchmarkGPT-5.4Claude Opus 4.6Diff
PinchBench v275.7017 / 45Reported best (effort unspecified)69.9026 / 45Reported best (effort unspecified)+5.80
Pinch Bench90.501 / 38Thinking (With Tools)87.408 / 38Thinking (With Tools)+3.10

General Knowledge

Claude Opus 4.6 2/2
BenchmarkGPT-5.4Claude Opus 4.6Diff
HLE11.30374 / 563Normal (No Tools) · Text only19.10299 / 563Normal (No Tools) · Text only-7.80
CritPt0.60168 / 200Normal (No Tools)2.80120 / 200Normal (No Tools)-2.20

Coding and Software Engineer

Claude Opus 4.6 1/1
BenchmarkGPT-5.4Claude Opus 4.6Diff
WeirdML v257.4431 / 52Normal (With Tools)65.9023 / 52Normal (With Tools)-8.46

General Evaluation

Claude Opus 4.6 1/1
BenchmarkGPT-5.4Claude Opus 4.6Diff
GPQA Diamond74.80284 / 462Normal (No Tools)84176 / 462Normal (No Tools)-9.20

Instruction Following

GPT-5.4 1/1
BenchmarkGPT-5.4Claude Opus 4.6Diff
IF Bench48.40164 / 282Normal (No Tools)44.60182 / 282Normal (No Tools)+3.80

Long Context

Claude Opus 4.6 1/1
BenchmarkGPT-5.4Claude Opus 4.6Diff
AA-LCR58.30130 / 170Normal (No Tools)67116 / 170Normal (No Tools)-8.70

Multimodal Understanding

Claude Opus 4.6 1/1
BenchmarkGPT-5.4Claude Opus 4.6Diff
MMMU-Pro70.60123 / 227Normal (No Tools)72.50114 / 227Normal (No Tools)-1.90

Text Embedding

Claude Opus 4.6 1/1
BenchmarkGPT-5.4Claude Opus 4.6Diff
Context Arena32.79112 / 126Normal (No Tools)60.3376 / 126Normal (No Tools)-27.54

Writing and Creative Capabilities

GPT-5.4 1/1
BenchmarkGPT-5.4Claude Opus 4.6Diff
Creative Writing1,83614 / 106Normal (No Tools)1,80419 / 106Normal (No Tools)+31.50

Specs

FieldGPT-5.4Claude Opus 4.6
PublisherOpenAIAnthropic
Release date2026-03-052026-02-05
Model typeMultimodal modelReasoning model
ArchitectureDenseDense
ParametersNot availableNot available
Context length1M1000K
Max output125K64K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGPT-5.4Claude Opus 4.6
Text input$2.5 / 1M tokens$5 / 1M tokens
Text output$15 / 1M tokens$25 / 1M tokens
Cache read$0.25 / 1M tokens$0.5 / 1M tokens
Cache writeNot public$6.25 / 1M tokens

Summary

  • GPT-5.4leads in:Claw-style Agent Evaluation (2/2), Instruction Following (1/1), Writing and Creative Capabilities (1/1)
  • Claude Opus 4.6leads in:Agent Level Benchmark (2/2), General Knowledge (2/2), Coding and Software Engineer (1/1), General Evaluation (1/1), Long Context (1/1), Multimodal Understanding (1/1), Text Embedding (1/1)

On average across the 13 shared benchmarks, Claude Opus 4.6 scores 4.05 higher.

Largest single-benchmark gap: Creative Writing — GPT-5.4 1,836 vs Claude Opus 4.6 1,804 (+31.50).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.