DataLearner logo

GPT-5.4vsGPT-5.2

Across 10 shared benchmarks, GPT-5.4 leads overall: GPT-5.4 wins 9, GPT-5.2 wins 1, with 0 ties and an average score difference of +18.76.

OpenAI
GPT-5.4

OpenAI · 2026-03-05 · Multimodal model

OpenAI
GPT-5.2

OpenAI · 2025-12-11 · Chat model

GPT-5.49 wins(90%)(10%)1 winGPT-5.2

Benchmark scores

Grouped by capability, sorted by largest gap within each. 10 shared benchmarks.

Coding and Software Engineer

Even 2/2
BenchmarkGPT-5.4GPT-5.2Diff
Text Arena (Coding)1,45724 / 35Thinking High (No Tools)1,48019 / 35Thinking High (No Tools)-22.77
WeirdML v257.4431 / 52Normal (With Tools)49.6037 / 52Normal (With Tools)+7.84

Math and Reasoning

GPT-5.4 2/2
BenchmarkGPT-5.4GPT-5.2Diff
FrontierMath Tier 4 v24912 / 42Extra-High (No Tools)31.7019 / 42Extra-High (No Tools)+17.30
FrontierMath v278.6010 / 58Extra-High (No Tools)67.4017 / 58Extra-High (No Tools)+11.20

Agent Level Benchmark

GPT-5.4 1/1
BenchmarkGPT-5.4GPT-5.2Diff
τ²-Bench - Telecom64.30151 / 264Normal (With Tools)46.50178 / 264Normal (With Tools)+17.80

AI Agent - Tool Usage

GPT-5.4 1/1
BenchmarkGPT-5.4GPT-5.2Diff
MCP-Atlas70.6028 / 43Extra-High (With Tools)67.6033 / 43Extra-High (With Tools)+3

General Evaluation

GPT-5.4 1/1
BenchmarkGPT-5.4GPT-5.2Diff
GPQA Diamond74.80284 / 462Normal (No Tools)73.23298 / 462Normal (No Tools)+1.57

General Knowledge

GPT-5.4 1/1
BenchmarkGPT-5.4GPT-5.2Diff
HLE11.30374 / 563Normal (No Tools) · Text only8419 / 563Normal (No Tools) · Text only+3.30

Long Context

GPT-5.4 1/1
BenchmarkGPT-5.4GPT-5.2Diff
AA-LCR58.30130 / 170Normal (No Tools)45.70144 / 170Normal (No Tools)+12.60

Writing and Creative Capabilities

GPT-5.4 1/1
BenchmarkGPT-5.4GPT-5.2Diff
Creative Writing1,83614 / 106Normal (No Tools)1,70026 / 106Normal (No Tools)+135.80

Specs

FieldGPT-5.4GPT-5.2
PublisherOpenAIOpenAI
Release date2026-03-052025-12-11
Model typeMultimodal modelChat model
ArchitectureDenseDense
ParametersNot availableNot available
Context length1M400K
Max output125KNot available

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGPT-5.4GPT-5.2
Text input$2.5 / 1M tokens$1.75 / 1M tokens
Text output$15 / 1M tokens$14 / 1M tokens
Cache read$0.25 / 1M tokens$0.175 / 1M tokens
Cache writeNot public$1.75 / 1M tokens

Summary

  • GPT-5.4leads in:Math and Reasoning (2/2), Agent Level Benchmark (1/1), AI Agent - Tool Usage (1/1), General Evaluation (1/1), General Knowledge (1/1), Long Context (1/1), Writing and Creative Capabilities (1/1)
  • Tied in:Coding and Software Engineer

On average across the 10 shared benchmarks, GPT-5.4 scores 18.76 higher.

Largest single-benchmark gap: Creative Writing — GPT-5.4 1,836 vs GPT-5.2 1,700 (+135.80).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.