DataLearner logo

GPT-5.1 Benchmark Details

GPT-5.1 currently shows benchmark results led by MMMU (2 / 28, score 85.40), Terminal Bench Hard (2 / 13, score 43), FrontierMath (13 / 60, score 26.70). This page also compares it with 2 competitor models and 2 predecessor or same-series models, including performance and pricing views when available. 2 source links are attached for reference.

Benchmark Results

GPT-5.1

Benchmark Results

Thinking
Tool usage
Internet

General Knowledge

13 evaluations
Benchmark / mode
Score
Rank/total
ARC-AGI
Thinking Level · Low
33.20
53 / 68
ARC-AGI
Thinking Level · Medium
57.70
40 / 68
ARC-AGI
Thinking Level · High
72.80
28 / 68
LiveBench
Standard Mode
42.65
106 / 115
LiveBench
Thinking Level · Low
59.95
71 / 115
LiveBench
Thinking Level · Medium
69.17
41 / 115
LiveBench
Thinking Level · High
72.04
29 / 115
HLE
Thinking Mode
26.50
108 / 185
HLE
Thinking Level · High
25.70
111 / 185
HLE
Thinking Level · HighToolsInternet
42.70
60 / 185
ARC-AGI-2
Thinking Level · Low
1.90
53 / 62
ARC-AGI-2
Thinking Level · Medium
6.50
44 / 62
ARC-AGI-2
Thinking Level · High
17.60
36 / 62

Other

2 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Mode
88.10
67 / 270
GPQA Diamond
Thinking Level · High
88.10
67 / 270

Coding and Software Engineer

8 evaluations
Benchmark / mode
Score
Rank/total
Text Arena (Coding)
Thinking Level · Low
1359
34 / 35
Text Arena (Coding)
Thinking Level · Medium
1387
32 / 35
SWE-bench Verified
Thinking Level · High
76.30
34 / 114
SWE-bench Verified
Thinking Level · HighTools
76.30
34 / 114
IC SWE-Lancer(Diamond)
Thinking Level · High
69.70
3 / 8
WeirdML v2
Thinking Level · HighTools
60.77
28 / 52
SWE-Bench Pro - Public
Thinking Level · High
50.80
45 / 59
GSO
Thinking Level · HighTools
13.70
13 / 21

Math and Reasoning

5 evaluations
Benchmark / mode
Score
Rank/total
AIME2025
Thinking Level · High
94
28 / 106
FrontierMath
Thinking Level · HighTools
26.70
13 / 60
FrontierMath - Tier 4
Thinking Level · Medium
4.20
40 / 80
FrontierMath - Tier 4
Thinking Level · High
12.50
29 / 80
FrontierMath - Tier 4
Thinking Level · HighTools
12.50
29 / 80

Multimodal Understanding

3 evaluations
Benchmark / mode
Score
Rank/total
MMMU
Thinking Level · High
85.40
2 / 28
VPCT
Thinking Level · Medium
53.30
9 / 24
VPCT
Thinking Level · High
58.70
7 / 24

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Thinking Level · High
53.20
28 / 67

Agent Level Benchmark

2 evaluations
Benchmark / mode
Score
Rank/total
τ²-Bench - Telecom
Thinking Level · HighTools
95.60
14 / 35
Terminal Bench Hard
Thinking Level · HighTools
43
2 / 13

AI Agent - Information Search

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
Thinking Level · High
50.80
43 / 54

AI Agent - Tool Usage

2 evaluations
Benchmark / mode
Score
Rank/total
MCP-Atlas
Thinking Level · HighTools
50.10
38 / 40
Terminal Bench 2.0
Thinking Level · HighTools
47.60
39 / 48

Competitor Comparison

Benchmark scores for GPT-5.1 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGPT-5.1CurrentClaude Opus 4Gemini 2.5-Pro
ARC-AGI
综合评估
72.80Thinking Level · High
35.70Standard Mode
37.00Thinking Enabled
ARC-AGI-2
综合评估
17.60Thinking Level · High
8.60Standard Mode
4.90Thinking Enabled
HLE
综合评估
42.70Thinking Level · High | Tools
10.70Standard Mode
21.60Thinking Enabled
LiveBench
综合评估
72.04Thinking Level · High
--
58.33Thinking Level · High
GPQA Diamond
科学与综合推理
88.10Thinking Level · High
79.60Standard Mode
86.40Thinking Enabled
GSO
编程与软件工程
13.70Thinking Level · High | Tools
--
3.92Standard Mode | Tools
SWE-bench Verified
编程与软件工程
76.30Thinking Level · High | Tools
72.50Standard Mode
67.20Thinking Enabled
AIME2025
数学推理
94.00Thinking Level · High
75.50Standard Mode
88.00Thinking Enabled
FrontierMath
数学推理
26.70Thinking Level · High | Tools
4.50Standard Mode
11.00Standard Mode
12.50Thinking Level · High | Tools
4.2032K
2.10Standard Mode
MMMU
多模态理解
85.40Thinking Level · High
--
82.00Thinking Enabled
VPCT
多模态理解
58.70Thinking Level · High
--
46.40Standard Mode
5 additional benchmarks remain in the chart above.

Standard API Pricing: GPT-5.1 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Gemini 2.5-Pro: Base price applies to <= 200000
ModelSupplierStandard inputStandard outputBase price applies to
GPT-5.1
OpenAI$1.25 / 1M tokens$10 / 1M tokens
Claude Opus 4
Anthropic$15 / 1M tokens$75 / 1M tokens
Gemini 2.5-Pro
Google Deep Mind$1.25 / 1M tokens$10 / 1M tokens<= 200000

Version History

How each version of the GPT-5.1 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGPT-5.1CurrentGPT-5GPT-4.5
ARC-AGI
综合评估
72.80Thinking Level · High
65.70Thinking Level · High
--
ARC-AGI-2
综合评估
17.60Thinking Level · High
9.90Thinking Level · High
--
HLE
综合评估
42.70Thinking Level · High | Tools
35.20Thinking Enabled | Tools
--
GPQA Diamond
科学与综合推理
88.10Thinking Level · High
87.30Thinking Enabled | Tools
71.40Standard Mode
GSO
编程与软件工程
13.70Thinking Level · High | Tools
6.90Thinking Level · High | Tools
--
IC SWE-Lancer(Diamond)
编程与软件工程
69.70Thinking Level · High
--
32.60Standard Mode
SWE-Bench Pro - Public
编程与软件工程
50.80Thinking Level · High
36.30Thinking Level · High
--
SWE-bench Verified
编程与软件工程
76.30Thinking Level · High | Tools
72.80Thinking Level · High
38.00Standard Mode
Text Arena (Coding)
编程与软件工程
1387.00Thinking Level · Medium
1395.00Thinking Level · Medium
--
WeirdML v2
编程与软件工程
60.77Thinking Level · High | Tools
60.70Thinking Level · High | Tools
--
AIME2025
数学推理
94.00Thinking Level · High
99.60Thinking Enabled | Tools
--
FrontierMath
数学推理
26.70Thinking Level · High | Tools
26.30Thinking Level · High | Tools
--
6 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: ARC-AGI · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the GPT-5.1 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
GPT-5.1
OpenAI$1.25 / 1M tokens$10 / 1M tokens
GPT-5
OpenAI$1.25 / 1M tokens$10 / 1M tokens

Sources