GPT-5.3 Codex Benchmark Details
GPT-5.3 Codex currently shows benchmark results led by Terminal Bench 2.0 (3 / 47, score 77.30), IC SWE-Lancer(Diamond) (1 / 8, score 81.40), WeirdML v2 (8 / 52, score 79.30). This page also compares it with 2 competitor models and 3 predecessor or same-series models, including performance and pricing views when available. 1 source link is attached for reference.
Benchmark Results
Benchmark Results
Coding and Software Engineer
5 evaluationsGeneral Knowledge
2 evaluationsAgent Level Benchmark
1 evaluationsCompetitor Comparison
Benchmark scores for GPT-5.3 Codex compared against top models in its class
4 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.
| Benchmark | GPT-5.3 CodexCurrent | Claude Opus 4.6 | Gemini 3.0 Pro (Preview 11-2025) |
|---|---|---|---|
Text Arena (Coding) 编程与软件工程 | 1406.65Standard Mode | 1555.35Standard Mode | -- |
WeirdML v2 编程与软件工程 | 79.30Thinking Level · Extra High | Tools | 77.95Thinking Level · High | Tools | -- |
LiveBench 综合评估 | 72.76Thinking Level · High | 76.33Thinking Level · High | 73.39Thinking Level · High |
Terminal Bench 2.0 AI Agent - 工具使用 | 77.30Thinking Level · Extra High | Tools | 65.40Extended Thinking | Tools | 56.90Thinking Level · High | Tools |
Standard API Pricing: GPT-5.3 Codex vs. Peer Models
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens
When a context threshold exists, the charted base price only applies within these limits:
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
GPT-5.3 Codex | OpenAI | $1.75 / 1M tokens | $14 / 1M tokens | — |
Claude Opus 4.6 | Anthropic | $5 / 1M tokens | $25 / 1M tokens | <= 200K |
Gemini 3.0 Pro (Preview 11-2025) | Google Deep Mind | $2 / 1M tokens | $12 / 1M tokens | <= 200000 |
Version History
How each version of the GPT-5.3 Codex series stacks up on benchmark tests
2 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.
| Benchmark | GPT-5.3 CodexCurrent | GPT-5.2-Codex | GPT-5.1-Codex-Max |
|---|---|---|---|
LiveBench 综合评估 | 72.76Thinking Level · High | 74.30Standard Mode | 73.98Deep Thinking Mode |
METR Time Horizons v1.1 Agent能力评测 | 349.53Standard Mode | Tools | -- | 161.75Standard Mode | Tools |
Single-Benchmark Version Trend
Viewing: LiveBench · 综合评估
Standard API Pricing Across the GPT-5.3 Codex Series
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
GPT-5.3 Codex | OpenAI | $1.75 / 1M tokens | $14 / 1M tokens | — |
GPT-5.2-Codex | OpenAI | $1.25 / 1M tokens | $10 / 1M tokens | — |
GPT-5.1-Codex-Max | OpenAI | $1.25 / 1M tokens | $10 / 1M tokens | — |
GPT-5 Codex | OpenAI | $1.25 / 1M tokens | $10 / 1M tokens | — |