DataLearner logo

GPT-5.3 Codex Benchmark Details

GPT-5.3 Codex currently shows benchmark results led by Terminal Bench 2.0 (3 / 48, score 77.30), Terminal Bench Hard (17 / 244, score 53), ECI (16 / 167, score 156.84). This page also compares it with 2 competitor models and 3 predecessor or same-series models, including performance and pricing views when available. 1 source link is attached for reference.

Benchmark Results

GPT-5.3 Codex

Benchmark Results

Thinking
Tool usage

Repository Engineering

2 evaluations
Benchmark / mode
Score
Rank/total
81.40
1 / 8
56.80
24 / 62

Cross-capability Suites

2 evaluations
Benchmark / mode
Score
Rank/total
72.76
24 / 117
71.64
31 / 117

Agentic Development

2 evaluations
Benchmark / mode
Score
Rank/total
Terminal Bench 2.0
Extra-HighTools
77.30
3 / 48
Terminal Bench Hard
Extra-HighTools
53
17 / 244

Visual Understanding

1 evaluations
Benchmark / mode
Score
Rank/total
MMMU-Pro
Extra-High
78.50
59 / 229

Scientific Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
CritPt
Extra-High
16.90
51 / 204

ML Engineering

2 evaluations
Benchmark / mode
Score
Rank/total
WeirdML v2
Standard ModeTools
79.30
8 / 52
WeirdML v2
Extra-HighTools
79.30
8 / 52

Preference Arenas

1 evaluations
Benchmark / mode
Score
Rank/total
Text Arena (Coding)
Standard Mode
1406.65
28 / 35

Capability Frontier Metrics

1 evaluations
Benchmark / mode
Score
Rank/total
METR Time Horizons v1.1
Standard ModeTools
349.53
4 / 22

Code Generation & Editing

1 evaluations
Benchmark / mode
Score
Rank/total
Vibe Code Bench v1.1
Extra-HighTools
61.77
31 / 60

Algorithmic Coding

1 evaluations
Benchmark / mode
Score
Rank/total
IOI (Vals v2)
Extra-HighTools
53.83
15 / 26

Capability Indices

1 evaluations
Benchmark / mode
Score
Rank/total
ECI
unknown
156.84
16 / 167

Competitor Comparison

Benchmark scores for GPT-5.3 Codex compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

9 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGPT-5.3 CodexCurrentClaude Opus 4.6Gemini 3.0 Pro (Preview 11-2025)
LiveBench
Accuracy
Cross-capability Suites
72.76Thinking Level · High
74.52Thinking Level · High
73.39Thinking Level · High
Terminal Bench 2.0
Accuracy
Agentic Development
77.30Thinking Level · Extra High | Tools
65.40Extended Thinking | Tools
56.90Thinking Level · High | Tools
Terminal Bench Hard
Accuracy
Agentic Development
53.00Thinking Level · Extra High | Tools
48.50Standard Mode | Tools
41.70Thinking Level · High | Tools
MMMU-Pro
Accuracy
Visual Understanding
78.50Thinking Level · Extra High
77.30Thinking Enabled | Tools
80.20Thinking Level · High
CritPt
Score
Scientific Reasoning
16.90Thinking Level · Extra High
12.60Thinking Level · High
9.10Thinking Level · High
WeirdML v2
Average accuracy across 17 tasks (%)
ML Engineering
79.30Thinking Level · Extra High | Tools
77.95Thinking Level · High | Tools
--
Text Arena (Coding)
Arena Score
Preference Arenas
1406.65Standard Mode
1555.35Standard Mode
--
Vibe Code Bench v1.1
Pass rate (%)
Code Generation & Editing
61.77Thinking Level · Extra High | Tools
57.57Standard Mode | Tools
--
ECI
ECI score (capability index, higher is better)
Capability Indices
156.84Thinking Level · High
155.36Thinking Level · High
--

Standard API Pricing: GPT-5.3 Codex vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Gemini 3.0 Pro (Preview 11-2025): Base price applies to <= 200000
ModelSupplierStandard inputStandard outputBase price applies to
GPT-5.3 Codex
OpenAI$1.75 / 1M tokens$14 / 1M tokens—
Claude Opus 4.6
Anthropic$5 / 1M tokens$25 / 1M tokens—
Gemini 3.0 Pro (Preview 11-2025)
Google DeepMind$2 / 1M tokens$12 / 1M tokens<= 200000

Version History

How each version of the GPT-5.3 Codex series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

6 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGPT-5.3 CodexCurrentGPT-5.2-CodexGPT-5.1-Codex-MaxGPT-5 Codex
LiveBench
Accuracy
Cross-capability Suites
72.76Thinking Level · High
73.98Standard Mode
73.98Deep Thinking Mode
--
Terminal Bench Hard
Accuracy
Agentic Development
53.00Thinking Level · Extra High | Tools
37.10Thinking Level · Extra High | Tools
--
37.90Thinking Level · High | Tools
MMMU-Pro
Accuracy
Visual Understanding
78.50Thinking Level · Extra High
76.30Thinking Level · Extra High
--
73.80Thinking Level · High
CritPt
Score
Scientific Reasoning
16.90Thinking Level · Extra High
8.70Thinking Level · Extra High
--
5.10Thinking Level · High
METR Time Horizons v1.1
50% task-completion time horizon (minutes)
Capability Frontier Metrics
349.53Standard Mode | Tools
--
161.75Standard Mode | Tools
--
Vibe Code Bench v1.1
Pass rate (%)
Code Generation & Editing
61.77Thinking Level · Extra High | Tools
37.91Thinking Level · High | Tools
22.17Thinking Level · High | Tools
--

Single-Benchmark Version Trend

Viewing: LiveBench · Cross-capability Suites

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the GPT-5.3 Codex Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
GPT-5.3 Codex
OpenAI$1.75 / 1M tokens$14 / 1M tokens—
GPT-5.2-Codex
OpenAI$1.25 / 1M tokens$10 / 1M tokens—
GPT-5.1-Codex-Max
OpenAI$1.25 / 1M tokens$10 / 1M tokens—
GPT-5 Codex
OpenAI$1.25 / 1M tokens$10 / 1M tokens—

Sources