GPT-5.3 CodexvsClaude Opus 4.6
GPT-5.3 Codex and Claude Opus 4.6 are tied across 4 shared benchmarks: GPT-5.3 Codex leads on 2, Claude Opus 4.6 leads on 2, with 0 ties and an average score difference of -31.74.
GPT-5.3 Codex
OpenAI · 2026-02-05 · Coding model
Claude Opus 4.6
Anthropic · 2026-02-05 · Reasoning model
GPT-5.3 Codex2 wins(50%)(50%)2 winsClaude Opus 4.6
Benchmark scores
Grouped by capability, sorted by largest gap within each. 4 shared benchmarks.
Coding and Software Engineer
Even 2/2| Benchmark | GPT-5.3 Codex | Claude Opus 4.6 | Diff |
|---|---|---|---|
| Text Arena (Coding) | 1,40728 / 35Normal (No Tools) | 1,5558 / 35Normal (No Tools) | -148.70 |
| WeirdML v2 | 79.308 / 52Normal (With Tools) | 65.9023 / 52Normal (With Tools) | +13.40 |
AI Agent - Tool Usage
GPT-5.3 Codex 1/1| Benchmark | GPT-5.3 Codex | Claude Opus 4.6 | Diff |
|---|---|---|---|
| Terminal Bench 2.0 | 77.303 / 47 | 65.4011 / 47Extended (with tools) | +11.90 |
General Knowledge
Claude Opus 4.6 1/1| Benchmark | GPT-5.3 Codex | Claude Opus 4.6 | Diff |
|---|---|---|---|
| LiveBench | 72.7625 / 115Thinking High (No Tools) | 76.338 / 115Thinking High (No Tools) | -3.57 |
Specs
| Field | GPT-5.3 Codex | Claude Opus 4.6 |
|---|---|---|
| Publisher | OpenAI | Anthropic |
| Release date | 2026-02-05 | 2026-02-05 |
| Model type | Coding model | Reasoning model |
| Architecture | Dense | Dense |
| Parameters | Not available | Not available |
| Context length | 400K | 1000K |
| Max output | 125K | 64K |
API pricing
Prices use DataLearner records when available; missing fields are not inferred.
| Item | GPT-5.3 Codex | Claude Opus 4.6 |
|---|---|---|
| Text input | $1.75 / 1M tokens | $0.5 / 1M tokens |
| Text output | $14 / 1M tokens | $25 / 1M tokens |
| Cache read | $0.175 / 1M tokens | $0.5 / 1M tokens |
| Cache write | Not public | $10 / 1M tokens |
Summary
- GPT-5.3 Codexleads in:AI Agent - Tool Usage (1/1)
- Claude Opus 4.6leads in:General Knowledge (1/1)
- Tied in:Coding and Software Engineer
On average across the 4 shared benchmarks, Claude Opus 4.6 scores 31.74 higher.
Largest single-benchmark gap: Text Arena (Coding) — GPT-5.3 Codex 1,407 vs Claude Opus 4.6 1,555 (-148.70).
Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.