DataLearner logo

GPT-5.1vsClaude Sonnet 4.5

Across 11 shared benchmarks, Claude Sonnet 4.5 leads overall: GPT-5.1 wins 3, Claude Sonnet 4.5 wins 8, with 0 ties and an average score difference of -6.37.

OpenAI
GPT-5.1

OpenAI · 2025-11-12 · Reasoning model

Anthropic
Claude Sonnet 4.5

Anthropic · 2025-09-30 · Chat model

GPT-5.13 wins(27%)(73%)8 winsClaude Sonnet 4.5

Benchmark scores

Grouped by capability, sorted by largest gap within each. 11 shared benchmarks.

Agent Level Benchmark

Claude Sonnet 4.5 2/2
BenchmarkGPT-5.1Claude Sonnet 4.5Diff
τ²-Bench - Telecom46.50178 / 264Normal (With Tools)70.50138 / 264Normal (With Tools)-24
Terminal Bench Hard22.70143 / 244Normal (With Tools)28.80116 / 244Normal (With Tools)-6.10

General Knowledge

Claude Sonnet 4.5 2/2
BenchmarkGPT-5.1Claude Sonnet 4.5Diff
LiveBench42.65108 / 117Normal (No Tools)53.6985 / 117Normal (No Tools)-11.04
HLE5.30471 / 563Normal (No Tools) · Text only7.20436 / 563Normal (No Tools) · Text only-1.90

Math and Reasoning

GPT-5.1 2/2
BenchmarkGPT-5.1Claude Sonnet 4.5Diff
IMO-ProofBench Advanced7.1016 / 24Thinking (No Tools)4.8019 / 24Thinking (No Tools)+2.30
AIME202538170 / 215Normal (No Tools)37173 / 215Normal (No Tools)+1

Coding and Software Engineer

Claude Sonnet 4.5 1/1
BenchmarkGPT-5.1Claude Sonnet 4.5Diff
LiveCodeBench49.40167 / 250Normal (No Tools)59136 / 250Normal (No Tools)-9.60

General Evaluation

Claude Sonnet 4.5 1/1
BenchmarkGPT-5.1Claude Sonnet 4.5Diff
GPQA Diamond64.30361 / 462Normal (No Tools)73.70292 / 462Normal (No Tools)-9.40

Instruction Following

GPT-5.1 1/1
BenchmarkGPT-5.1Claude Sonnet 4.5Diff
IF Bench43.20192 / 282Normal (No Tools)42.70199 / 282Normal (No Tools)+0.50

Long Context

Claude Sonnet 4.5 1/1
BenchmarkGPT-5.1Claude Sonnet 4.5Diff
AA-LCR45147 / 170Normal (No Tools)54133 / 170Normal (No Tools)-9

Multimodal Understanding

Claude Sonnet 4.5 1/1
BenchmarkGPT-5.1Claude Sonnet 4.5Diff
MMMU-Pro62.40163 / 227Normal (No Tools)65.20151 / 227Normal (No Tools)-2.80

Specs

FieldGPT-5.1Claude Sonnet 4.5
PublisherOpenAIAnthropic
Release date2025-11-122025-09-30
Model typeReasoning modelChat model
ArchitectureDenseDense
ParametersNot availableNot available
Context length400K1000K
Max output128K64K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGPT-5.1Claude Sonnet 4.5
Text input$1.25 / 1M tokens$3 / 1M tokens
Text output$10 / 1M tokens$15 / 1M tokens
Cache read$0.125 / 1M tokens$0.3 / 1M tokens
Cache write$0 / 1M tokens$3.75 / 1M tokens

Summary

  • GPT-5.1leads in:Math and Reasoning (2/2), Instruction Following (1/1)
  • Claude Sonnet 4.5leads in:Agent Level Benchmark (2/2), General Knowledge (2/2), Coding and Software Engineer (1/1), General Evaluation (1/1), Long Context (1/1), Multimodal Understanding (1/1)

On average across the 11 shared benchmarks, Claude Sonnet 4.5 scores 6.37 higher.

Largest single-benchmark gap: τ²-Bench - Telecom — GPT-5.1 46.50 vs Claude Sonnet 4.5 70.50 (-24).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.