GPT-5.4vsClaude Opus 4.6
Across 13 shared benchmarks, Claude Opus 4.6 leads overall: GPT-5.4 wins 4, Claude Opus 4.6 wins 9, with 0 ties and an average score difference of -4.05.
GPT-5.4
OpenAI · 2026-03-05 · Multimodal model
Claude Opus 4.6
Anthropic · 2026-02-05 · Reasoning model
GPT-5.44 wins(31%)(69%)9 winsClaude Opus 4.6
Benchmark scores
Grouped by capability, sorted by largest gap within each. 13 shared benchmarks.
Agent Level Benchmark
Claude Opus 4.6 2/2| Benchmark | GPT-5.4 | Claude Opus 4.6 | Diff |
|---|---|---|---|
| τ²-Bench - Telecom | 64.30151 / 264Normal (With Tools) | 84.8090 / 264Normal (With Tools) | -20.50 |
| Terminal Bench Hard | 37.9059 / 244Normal (With Tools) | 48.5025 / 244Normal (With Tools) | -10.60 |
Claw-style Agent Evaluation
GPT-5.4 2/2| Benchmark | GPT-5.4 | Claude Opus 4.6 | Diff |
|---|---|---|---|
| PinchBench v2 | 75.7017 / 45Reported best (effort unspecified) | 69.9026 / 45Reported best (effort unspecified) | +5.80 |
| Pinch Bench | 90.501 / 38Thinking (With Tools) | 87.408 / 38Thinking (With Tools) | +3.10 |
General Knowledge
Claude Opus 4.6 2/2| Benchmark | GPT-5.4 | Claude Opus 4.6 | Diff |
|---|---|---|---|
| HLE | 11.30374 / 563Normal (No Tools) · Text only | 19.10299 / 563Normal (No Tools) · Text only | -7.80 |
| CritPt | 0.60168 / 200Normal (No Tools) | 2.80120 / 200Normal (No Tools) | -2.20 |
Coding and Software Engineer
Claude Opus 4.6 1/1| Benchmark | GPT-5.4 | Claude Opus 4.6 | Diff |
|---|---|---|---|
| WeirdML v2 | 57.4431 / 52Normal (With Tools) | 65.9023 / 52Normal (With Tools) | -8.46 |
General Evaluation
Claude Opus 4.6 1/1| Benchmark | GPT-5.4 | Claude Opus 4.6 | Diff |
|---|---|---|---|
| GPQA Diamond | 74.80284 / 462Normal (No Tools) | 84176 / 462Normal (No Tools) | -9.20 |
Instruction Following
GPT-5.4 1/1| Benchmark | GPT-5.4 | Claude Opus 4.6 | Diff |
|---|---|---|---|
| IF Bench | 48.40164 / 282Normal (No Tools) | 44.60182 / 282Normal (No Tools) | +3.80 |
Long Context
Claude Opus 4.6 1/1| Benchmark | GPT-5.4 | Claude Opus 4.6 | Diff |
|---|---|---|---|
| AA-LCR | 58.30130 / 170Normal (No Tools) | 67116 / 170Normal (No Tools) | -8.70 |
Multimodal Understanding
Claude Opus 4.6 1/1| Benchmark | GPT-5.4 | Claude Opus 4.6 | Diff |
|---|---|---|---|
| MMMU-Pro | 70.60123 / 227Normal (No Tools) | 72.50114 / 227Normal (No Tools) | -1.90 |
Text Embedding
Claude Opus 4.6 1/1| Benchmark | GPT-5.4 | Claude Opus 4.6 | Diff |
|---|---|---|---|
| Context Arena | 32.79112 / 126Normal (No Tools) | 60.3376 / 126Normal (No Tools) | -27.54 |
Writing and Creative Capabilities
GPT-5.4 1/1| Benchmark | GPT-5.4 | Claude Opus 4.6 | Diff |
|---|---|---|---|
| Creative Writing | 1,83614 / 106Normal (No Tools) | 1,80419 / 106Normal (No Tools) | +31.50 |
Specs
| Field | GPT-5.4 | Claude Opus 4.6 |
|---|---|---|
| Publisher | OpenAI | Anthropic |
| Release date | 2026-03-05 | 2026-02-05 |
| Model type | Multimodal model | Reasoning model |
| Architecture | Dense | Dense |
| Parameters | Not available | Not available |
| Context length | 1M | 1000K |
| Max output | 125K | 64K |
API pricing
Prices use DataLearner records when available; missing fields are not inferred.
| Item | GPT-5.4 | Claude Opus 4.6 |
|---|---|---|
| Text input | $2.5 / 1M tokens | $5 / 1M tokens |
| Text output | $15 / 1M tokens | $25 / 1M tokens |
| Cache read | $0.25 / 1M tokens | $0.5 / 1M tokens |
| Cache write | Not public | $6.25 / 1M tokens |
Summary
- GPT-5.4leads in:Claw-style Agent Evaluation (2/2), Instruction Following (1/1), Writing and Creative Capabilities (1/1)
- Claude Opus 4.6leads in:Agent Level Benchmark (2/2), General Knowledge (2/2), Coding and Software Engineer (1/1), General Evaluation (1/1), Long Context (1/1), Multimodal Understanding (1/1), Text Embedding (1/1)
On average across the 13 shared benchmarks, Claude Opus 4.6 scores 4.05 higher.
Largest single-benchmark gap: Creative Writing — GPT-5.4 1,836 vs Claude Opus 4.6 1,804 (+31.50).
Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.