DataLearner logo
GP

GPT-5.1

Reasoning modelGPTGPT-5.1

GPT-5.1

Release date: 2025-11-12Updated: 2026-06-15Views: 1,035
Live demoGitHubHugging FaceCompare
Parameters
No data
Context length
400K
Multilingual
Supported
Reasoning ability
4/5

GPT-5.1 is an AI model published by OpenAI, released on 2025-11-12, for Reasoning model, and 400K context length, with a 1387.00 score on Text Arena (Coding).

Data sourced primarily from official releases (GitHub, Hugging Face, papers), then benchmark leaderboards, then third-party evaluators. Learn about our data methodology

GPT-5.1

Model basics

Reasoning traces
Supported
Thinking modes
Standard Mode (Default)Thinking Mode
Context length
400K tokens
Max output length
128K tokens
Model type
Reasoning model
Modality (in / out)
Text, Image → Text
Release date
2025-11-12
Model file size
No data
MoE architecture
No
Total params / Active params
No data / Not applicable
Knowledge cutoff
No data
GPT-5.1

Open source & experience

Code license
Proprietary
Weights license
Proprietary
GitHub repo
N/A
Hugging Face
N/A
Live demo
GPT-5.1

Official resources

Paper
DataLearnerAI blog
GPT-5.1

API details

API speed
3/5
💡Default unit: $/1M tokens. If vendors use other units, follow their published pricing.
Standard
TypeConditionInputOutput
Text-$1.25/ 1M tokens$10.00/ 1M tokens
Image-$1.25/ 1M tokens
Batch
TypeConditionInputOutput
Text-$0.625/ 1M tokens$5.00/ 1M tokens
Image-$0.625/ 1M tokens
Cache PricingPrompt Cache
TypeTTLWriteRead
Text-$0.0000/ 1M tokens
Cache = write
Text-$0.0000/ 1M tokens$0.125/ 1M tokens
Text-$0.125/ 1M tokens
Cache = hit
Image-$0.0000/ 1M tokens
Cache = write
Image-$0.125/ 1M tokens
Cache = hit

“—” means the modality is not billed in that direction, or the vendor has not published a price for it.

GPT-5.1

Benchmark Results

GPT-5.1 currently shows benchmark results led by MMMU (2 / 28, score 85.40), Terminal Bench Hard (2 / 13, score 43), FrontierMath (13 / 60, score 26.70). This page also consolidates core specs, context limits, and API pricing so you can evaluate the model from benchmark results and deployment constraints together.

Thinking
Tool usage
Internet

General Knowledge

13 evaluations
Benchmark / mode
Score
Rank/total
33.20
77 / 92
ARC-AGI-1
Medium
57.70
64 / 92
72.80
52 / 92
LiveBench
Standard Mode
42.65
108 / 117
59.95
73 / 117
LiveBench
Medium
69.17
40 / 117
72.04
29 / 117
HLE
Thinking Mode
26.50
118 / 197
HLE
High
25.70
121 / 197
HLE
HighToolsInternet
42.70
66 / 197
1.90
76 / 85
ARC-AGI-2
Medium
6.50
67 / 85
17.60
59 / 85

General Evaluation

2 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Mode
88.10
70 / 274
88.10
70 / 274

Coding and Software Engineer

8 evaluations
Benchmark / mode
Score
Rank/total
1359
34 / 35
1387
32 / 35
76.30
34 / 116
76.30
34 / 116
WeirdML v2
HighTools
60.77
28 / 52
50.80
46 / 62
GSO
HighTools
13.70
13 / 21

Math and Reasoning

5 evaluations
Benchmark / mode
Score
Rank/total
94
28 / 107
FrontierMath
HighTools
26.70
13 / 60
4.20
40 / 80
12.50
29 / 80
12.50
29 / 80

Multimodal Understanding

3 evaluations
Benchmark / mode
Score
Rank/total
MMMU
High
85.40
2 / 28
VPCT
Medium
53.30
9 / 24
VPCT
High
58.70
7 / 24

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
53.20
45 / 94

Agent Level Benchmark

2 evaluations
Benchmark / mode
Score
Rank/total
95.60
15 / 36
43
2 / 13

AI Agent - Information Search

1 evaluations
Benchmark / mode
Score
Rank/total
50.80
45 / 57

AI Agent - Tool Usage

2 evaluations
Benchmark / mode
Score
Rank/total
MCP-Atlas
HighTools
50.10
39 / 41
47.60
39 / 48

Compare with other models

GPT-5.1

Publisher

GPT-5.1

Model Overview

GPT-5.1 is an AI model published by OpenAI, released on 2025-11-12, for Reasoning model, and 400K context length, with a 1387.00 score on Text Arena (Coding).

DataLearner on WeChat

Follow DataLearner on WeChat for AI model updates and research notes.

DataLearner WeChat QR code