DataLearner 标志

GPT-5.1vsGPT-5

在 23 个共同 benchmark 中,GPT-5 整体领先:GPT-5.1 领先 8 项,GPT-5 领先 13 项,持平 2 项,平均分差 -6.28。

OpenAI
GPT-5.1

OpenAI · 2025-11-12 · 推理大模型

OpenAI
GPT-5

OpenAI · 2025-08-07 · 基础大模型

GPT-5.18 (35%)持平2(57%)13 GPT-5

评测分数

按能力类目分组,每组内按分差大小排列;共 23 项。

Coding and Software Engineer

GPT-5.1 领先 4/6
评测项GPT-5.1GPT-5分差
SWE-Bench Pro - Public50.8046 / 62Thinking High (No Tools)36.3059 / 62Thinking High (No Tools)+14.50
Text Arena (Coding)1,38732 / 35Thinking Medium (No Tools)1,39529 / 35Thinking Medium (No Tools)-8
GSO13.7013 / 21Thinking High (With Tools)6.9016 / 21Thinking High (With Tools)+6.80
LiveCodeBench49.40167 / 250Normal (No Tools)55.80149 / 250Normal (No Tools)-6.40
SWE-bench Verified76.3034 / 116Thinking High (No Tools)72.8051 / 116Thinking High (No Tools)+3.50
WeirdML v260.7728 / 52Thinking High (With Tools)60.7029 / 52Thinking High (With Tools)+0.07

Math and Reasoning

GPT-5 领先 2/5
评测项GPT-5.1GPT-5分差
AIME202538170 / 215Normal (No Tools)61.90135 / 215Normal (No Tools)-23.90
IMO-ProofBench Advanced7.1016 / 24Thinking (No Tools)2011 / 24Thinking (No Tools)-12.90
FrontierMath26.7013 / 60Thinking High (With Tools)26.3014 / 60Thinking High (With Tools)+0.40
FrontierMath - Tier 412.5029 / 80Thinking High (No Tools)12.5029 / 80Thinking High (No Tools)持平
MathArena Apex1.0414 / 17Thinking High (With Tools)1.0414 / 17Thinking High (With Tools)持平

Agent Level Benchmark

GPT-5 领先 2/3
评测项GPT-5.1GPT-5分差
τ²-Bench - Telecom46.50178 / 264Normal (With Tools)67144 / 264Normal (With Tools)-20.50
τ³-Banking15.90103 / 164Thinking High (With Tools)22.1085 / 164Thinking High (With Tools)-6.20
Terminal Bench Hard22.70143 / 244Normal (With Tools)18.20154 / 244Normal (With Tools)+4.50

General Knowledge

GPT-5 领先 2/2
评测项GPT-5.1GPT-5分差
HLE5.30471 / 563Normal (No Tools) · Text only6.60450 / 563Normal (No Tools) · Text only-1.30
CritPt4.9098 / 200Thinking High (No Tools)5.7088 / 200Thinking High (No Tools)-0.80

Multimodal Understanding

胶着 2/2
评测项GPT-5.1GPT-5分差
VPCT58.707 / 24Thinking High (No Tools)665 / 24Thinking High (No Tools)-7.30
MMMU-Pro62.40163 / 227Normal (No Tools)62.10165 / 227Normal (No Tools)+0.30

AI Agent - Tool Usage

GPT-5.1 领先 1/1
评测项GPT-5.1GPT-5分差
Terminal-Bench 2.152.40126 / 192Thinking High (With Tools)35.20151 / 192Thinking High (With Tools)+17.20

Instruction Following

GPT-5 领先 1/1
评测项GPT-5.1GPT-5分差
IF Bench43.20192 / 282Normal (No Tools)45.60178 / 282Normal (No Tools)-2.40

Productivity Knowledge

GPT-5 领先 1/1
评测项GPT-5.1GPT-5分差
GDPval-AA v293093 / 105Thinking High (With Tools)1,01589 / 105Thinking High (With Tools)-85

常识推理

GPT-5 领先 1/1
评测项GPT-5.1GPT-5分差
SimpleBench53.2045 / 93Thinking High (No Tools)56.7040 / 93Thinking High (No Tools)-3.50

科学与综合推理

GPT-5 领先 1/1
评测项GPT-5.1GPT-5分差
GPQA Diamond64.30361 / 462Normal (No Tools)77.80257 / 462Normal (No Tools)-13.50

规格对比

字段GPT-5.1GPT-5
发布机构OpenAIOpenAI
发布时间2025-11-122025-08-07
模型类型推理大模型基础大模型
架构稠密模型稠密模型
参数规模暂无数据暂无数据
上下文长度400K400K
最大输出128K128K

API 调用价格

价格优先使用 DataLearner 配置的 API 记录;缺失项不做推测。

价格项GPT-5.1GPT-5
文本输入$1.25 / 1M tokens$1.25 / 1M tokens
文本输出$10 / 1M tokens$10 / 1M tokens
缓存读取$0.125 / 1M tokens$0.125 / 1M tokens
缓存写入$0 / 1M tokens$0 / 1M tokens

小结

  • GPT-5.1在以下类目领先:Coding and Software Engineer (4/6)、AI Agent - Tool Usage (1/1)
  • GPT-5在以下类目领先:Math and Reasoning (2/5)、Agent Level Benchmark (2/3)、General Knowledge (2/2)、Instruction Following (1/1)、Productivity Knowledge (1/1)、常识推理 (1/1)、科学与综合推理 (1/1)
  • 胶着类目:Multimodal Understanding

23 个共同 benchmark 上,GPT-5 平均高出 6.28 分。

单项差距最大的 benchmark:GDPval-AA v2 — GPT-5.1 930,GPT-5 1,015(分差 -85)。

本页正文由结构化模型、价格与 benchmark 数据生成,不使用实时 LLM 撰写。