DataLearner logo

GPT-5.3 CodexvsGPT-5.2-Codex

Across 8 shared benchmarks, GPT-5.3 Codex leads overall: GPT-5.3 Codex wins 6, GPT-5.2-Codex wins 2, with 0 ties and an average score difference of +3.43.

OpenAI
GPT-5.3 Codex

OpenAI · 2026-02-05 · Coding model

OpenAI
GPT-5.2-Codex

OpenAI · 2025-12-18 · Coding model

GPT-5.3 Codex6 wins(75%)(25%)2 winsGPT-5.2-Codex

Benchmark scores

Grouped by capability, sorted by largest gap within each. 8 shared benchmarks.

Agent Level Benchmark

Even 2/2
BenchmarkGPT-5.3 CodexGPT-5.2-CodexDiff
Terminal Bench Hard5317 / 244Extra-High (With Tools)37.1065 / 244Extra-High (With Tools)+15.90
τ²-Bench - Telecom8685 / 264Extra-High (With Tools)92.1054 / 264Extra-High (With Tools)-6.10

General Knowledge

GPT-5.3 Codex 2/2
BenchmarkGPT-5.3 CodexGPT-5.2-CodexDiff
CritPt16.9047 / 200Extra-High (No Tools)8.7078 / 200Extra-High (No Tools)+8.20
HLE42.50109 / 563Extra-High (No Tools) · Text only35.70164 / 563Extra-High (No Tools) · Text only+6.80

General Evaluation

GPT-5.3 Codex 1/1
BenchmarkGPT-5.3 CodexGPT-5.2-CodexDiff
GPQA Diamond91.5061 / 462Extra-High (No Tools)89.9086 / 462Extra-High (No Tools)+1.60

Instruction Following

GPT-5.2-Codex 1/1
BenchmarkGPT-5.3 CodexGPT-5.2-CodexDiff
IF Bench75.4034 / 282Extra-High (No Tools)77.6016 / 282Extra-High (No Tools)-2.20

Long Context

GPT-5.3 Codex 1/1
BenchmarkGPT-5.3 CodexGPT-5.2-CodexDiff
AA-LCR83.3011 / 170Extra-High (No Tools)82.3020 / 170Extra-High (No Tools)+1

Multimodal Understanding

GPT-5.3 Codex 1/1
BenchmarkGPT-5.3 CodexGPT-5.2-CodexDiff
MMMU-Pro78.5059 / 227Extra-High (No Tools)76.3075 / 227Extra-High (No Tools)+2.20

Specs

FieldGPT-5.3 CodexGPT-5.2-Codex
PublisherOpenAIOpenAI
Release date2026-02-052025-12-18
Model typeCoding modelCoding model
ArchitectureDenseDense
ParametersNot availableNot available
Context length400KNot available
Max output125KNot available

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGPT-5.3 CodexGPT-5.2-Codex
Text input$1.75 / 1M tokens$1.25 / 1M tokens
Text output$14 / 1M tokens$10 / 1M tokens
Cache read$0.175 / 1M tokens$0.125 / 1M tokens
Cache writeNot public$0 / 1M tokens

Summary

  • GPT-5.3 Codexleads in:General Knowledge (2/2), General Evaluation (1/1), Long Context (1/1), Multimodal Understanding (1/1)
  • GPT-5.2-Codexleads in:Instruction Following (1/1)
  • Tied in:Agent Level Benchmark

On average across the 8 shared benchmarks, GPT-5.3 Codex scores 3.43 higher.

Largest single-benchmark gap: Terminal Bench Hard — GPT-5.3 Codex 53 vs GPT-5.2-Codex 37.10 (+15.90).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.