DataLearner logo

GPT-5.2 ProvsOpus 4.5

Across 6 shared benchmarks, GPT-5.2 Pro leads overall: GPT-5.2 Pro wins 5, Opus 4.5 wins 1, with 0 ties and an average score difference of +10.43.

OpenAI
GPT-5.2 Pro

OpenAI · 2025-12-11 · Reasoning model

Anthropic
Opus 4.5

Anthropic · 2025-11-25 · Reasoning model

GPT-5.2 Pro5 wins(83%)(17%)1 winOpus 4.5

Benchmark scores

Grouped by capability, sorted by largest gap within each. 6 shared benchmarks.

General Knowledge

GPT-5.2 Pro 4/4
BenchmarkGPT-5.2 ProOpus 4.5Diff
ARC-AGI-254.2023 / 6237.6029 / 62Extended (no tools)+16.60
ARC-AGI90.5017 / 688024 / 68Extended (no tools)+10.50
HLE5029 / 17243.2049 / 172Extended (with tools)+6.80
GPQA Diamond93.209 / 1878742 / 187Extended (no tools)+6.20

Commonsense Reasoning

Opus 4.5 1/1
BenchmarkGPT-5.2 ProOpus 4.5Diff
Simple Bench57.4019 / 63极高强度思考(无工具)6212 / 63Extended (no tools)-4.60

Math and Reasoning

GPT-5.2 Pro 1/1
BenchmarkGPT-5.2 ProOpus 4.5Diff
FrontierMath - Tier 431.309 / 804.2040 / 80Normal (No Tools)+27.10

Specs

FieldGPT-5.2 ProOpus 4.5
PublisherOpenAIAnthropic
Release date2025-12-112025-11-25
Model typeReasoning modelReasoning model
ArchitectureDenseDense
ParametersNot availableNot available
Context length256K200K
Max outputNot available64K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGPT-5.2 ProOpus 4.5
Text input$21 / 1M tokens$5 / 1M tokens
Text output$168 / 1M tokens$25 / 1M tokens
Cache readNot public$0.5 / 1M tokens
Cache writeNot public$6.25 / 1M tokens

Summary

  • GPT-5.2 Proleads in:General Knowledge (4/4), Math and Reasoning (1/1)
  • Opus 4.5leads in:Commonsense Reasoning (1/1)

On average across the 6 shared benchmarks, GPT-5.2 Pro scores 10.43 higher.

Largest single-benchmark gap: FrontierMath - Tier 4 — GPT-5.2 Pro 31.30 vs Opus 4.5 4.20 (+27.10).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.