DataLearner logo

GPT-5.1vsGPT-5

Across 23 shared benchmarks, GPT-5 leads overall: GPT-5.1 wins 8, GPT-5 wins 13, with 2 ties and an average score difference of -6.28.

OpenAI
GPT-5.1

OpenAI · 2025-11-12 · Reasoning model

OpenAI
GPT-5

OpenAI · 2025-08-07 · Foundation model

GPT-5.18 wins(35%)Ties2(57%)13 winsGPT-5

Benchmark scores

Grouped by capability, sorted by largest gap within each. 23 shared benchmarks.

Coding and Software Engineer

GPT-5.1 4/6
BenchmarkGPT-5.1GPT-5Diff
SWE-Bench Pro - Public50.8046 / 62Thinking High (No Tools)36.3059 / 62Thinking High (No Tools)+14.50
Text Arena (Coding)1,38732 / 35Thinking Medium (No Tools)1,39529 / 35Thinking Medium (No Tools)-8
GSO13.7013 / 21Thinking High (With Tools)6.9016 / 21Thinking High (With Tools)+6.80
LiveCodeBench49.40167 / 250Normal (No Tools)55.80149 / 250Normal (No Tools)-6.40
SWE-bench Verified76.3034 / 116Thinking High (No Tools)72.8051 / 116Thinking High (No Tools)+3.50
WeirdML v260.7728 / 52Thinking High (With Tools)60.7029 / 52Thinking High (With Tools)+0.07

Math and Reasoning

GPT-5 2/5
BenchmarkGPT-5.1GPT-5Diff
AIME202538170 / 215Normal (No Tools)61.90135 / 215Normal (No Tools)-23.90
IMO-ProofBench Advanced7.1016 / 24Thinking (No Tools)2011 / 24Thinking (No Tools)-12.90
FrontierMath26.7013 / 60Thinking High (With Tools)26.3014 / 60Thinking High (With Tools)+0.40
FrontierMath - Tier 412.5029 / 80Thinking High (No Tools)12.5029 / 80Thinking High (No Tools)
MathArena Apex1.0414 / 17Thinking High (With Tools)1.0414 / 17Thinking High (With Tools)

Agent Level Benchmark

GPT-5 2/3
BenchmarkGPT-5.1GPT-5Diff
τ²-Bench - Telecom46.50178 / 264Normal (With Tools)67144 / 264Normal (With Tools)-20.50
τ³-Banking15.90103 / 164Thinking High (With Tools)22.1085 / 164Thinking High (With Tools)-6.20
Terminal Bench Hard22.70143 / 244Normal (With Tools)18.20154 / 244Normal (With Tools)+4.50

General Knowledge

GPT-5 2/2
BenchmarkGPT-5.1GPT-5Diff
HLE5.30471 / 563Normal (No Tools) · Text only6.60450 / 563Normal (No Tools) · Text only-1.30
CritPt4.9098 / 200Thinking High (No Tools)5.7088 / 200Thinking High (No Tools)-0.80

Multimodal Understanding

Even 2/2
BenchmarkGPT-5.1GPT-5Diff
VPCT58.707 / 24Thinking High (No Tools)665 / 24Thinking High (No Tools)-7.30
MMMU-Pro62.40163 / 227Normal (No Tools)62.10165 / 227Normal (No Tools)+0.30

AI Agent - Tool Usage

GPT-5.1 1/1
BenchmarkGPT-5.1GPT-5Diff
Terminal-Bench 2.152.40126 / 192Thinking High (With Tools)35.20151 / 192Thinking High (With Tools)+17.20

Commonsense Reasoning

GPT-5 1/1
BenchmarkGPT-5.1GPT-5Diff
SimpleBench53.2045 / 93Thinking High (No Tools)56.7040 / 93Thinking High (No Tools)-3.50

General Evaluation

GPT-5 1/1
BenchmarkGPT-5.1GPT-5Diff
GPQA Diamond64.30361 / 462Normal (No Tools)77.80257 / 462Normal (No Tools)-13.50

Instruction Following

GPT-5 1/1
BenchmarkGPT-5.1GPT-5Diff
IF Bench43.20192 / 282Normal (No Tools)45.60178 / 282Normal (No Tools)-2.40

Productivity Knowledge

GPT-5 1/1
BenchmarkGPT-5.1GPT-5Diff
GDPval-AA v293093 / 105Thinking High (With Tools)1,01589 / 105Thinking High (With Tools)-85

Specs

FieldGPT-5.1GPT-5
PublisherOpenAIOpenAI
Release date2025-11-122025-08-07
Model typeReasoning modelFoundation model
ArchitectureDenseDense
ParametersNot availableNot available
Context length400K400K
Max output128K128K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGPT-5.1GPT-5
Text input$1.25 / 1M tokens$1.25 / 1M tokens
Text output$10 / 1M tokens$10 / 1M tokens
Cache read$0.125 / 1M tokens$0.125 / 1M tokens
Cache write$0 / 1M tokens$0 / 1M tokens

Summary

  • GPT-5.1leads in:Coding and Software Engineer (4/6), AI Agent - Tool Usage (1/1)
  • GPT-5leads in:Math and Reasoning (2/5), Agent Level Benchmark (2/3), General Knowledge (2/2), Commonsense Reasoning (1/1), General Evaluation (1/1), Instruction Following (1/1), Productivity Knowledge (1/1)
  • Tied in:Multimodal Understanding

On average across the 23 shared benchmarks, GPT-5 scores 6.28 higher.

Largest single-benchmark gap: GDPval-AA v2 — GPT-5.1 930 vs GPT-5 1,015 (-85).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.