DataLearner logo

GPT-5.3 Codex Benchmark Details

GPT-5.3 Codex currently shows benchmark results led by Terminal Bench 2.0 (3 / 47, score 77.30), IC SWE-Lancer(Diamond) (1 / 8, score 81.40), WeirdML v2 (8 / 52, score 79.30). This page also compares it with 2 competitor models and 3 predecessor or same-series models, including performance and pricing views when available. 1 source link is attached for reference.

Benchmark Results

GPT-5.3 Codex

Benchmark Results

Thinking
Tool usage

Coding and Software Engineer

5 evaluations
Benchmark / mode
Score
Rank/total
Text Arena (Coding)
Standard Mode
1406.65
28 / 35
WeirdML v2
Standard ModeTools
79.30
8 / 52
WeirdML v2
Extra-HighTools
79.30
8 / 52

General Knowledge

2 evaluations
Benchmark / mode
Score
Rank/total
72.76
25 / 115
LiveBench
Deep Thinking Mode
71.64
32 / 115

AI Agent - Tool Usage

1 evaluations
Benchmark / mode
Score
Rank/total

Agent Level Benchmark

1 evaluations
Benchmark / mode
Score
Rank/total
METR Time Horizons v1.1
Standard ModeTools
349.53
4 / 22

Competitor Comparison

Benchmark scores for GPT-5.3 Codex compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

4 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGPT-5.3 CodexCurrentClaude Opus 4.6Gemini 3.0 Pro (Preview 11-2025)
Text Arena (Coding)
编程与软件工程
1406.65Standard Mode
1555.35Standard Mode
--
WeirdML v2
编程与软件工程
79.30Thinking Level · Extra High | Tools
77.95Thinking Level · High | Tools
--
LiveBench
综合评估
72.76Thinking Level · High
76.33Thinking Level · High
73.39Thinking Level · High
Terminal Bench 2.0
AI Agent - 工具使用
77.30Thinking Level · Extra High | Tools
65.40Extended Thinking | Tools
56.90Thinking Level · High | Tools

Standard API Pricing: GPT-5.3 Codex vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Claude Opus 4.6: Base price applies to <= 200K
Gemini 3.0 Pro (Preview 11-2025): Base price applies to <= 200000
ModelSupplierStandard inputStandard outputBase price applies to
GPT-5.3 Codex
OpenAI$1.75 / 1M tokens$14 / 1M tokens
Claude Opus 4.6
Anthropic$5 / 1M tokens$25 / 1M tokens<= 200K
Gemini 3.0 Pro (Preview 11-2025)
Google Deep Mind$2 / 1M tokens$12 / 1M tokens<= 200000

Version History

How each version of the GPT-5.3 Codex series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

2 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGPT-5.3 CodexCurrentGPT-5.2-CodexGPT-5.1-Codex-Max
LiveBench
综合评估
72.76Thinking Level · High
74.30Standard Mode
73.98Deep Thinking Mode
METR Time Horizons v1.1
Agent能力评测
349.53Standard Mode | Tools
--
161.75Standard Mode | Tools

Single-Benchmark Version Trend

Viewing: LiveBench · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the GPT-5.3 Codex Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
GPT-5.3 Codex
OpenAI$1.75 / 1M tokens$14 / 1M tokens
GPT-5.2-Codex
OpenAI$1.25 / 1M tokens$10 / 1M tokens
GPT-5.1-Codex-Max
OpenAI$1.25 / 1M tokens$10 / 1M tokens
GPT-5 Codex
OpenAI$1.25 / 1M tokens$10 / 1M tokens

Sources