DataLearner logo

GPT-5 Benchmark Details

GPT-5 currently shows benchmark results led by Aider-Polyglot (1 / 59, score 88), AIME2025 (10 / 215, score 99.60), Fiction.liveBench (1 / 16, score 97.20). This page also compares it with 2 competitor models and 3 predecessor or same-series models, including performance and pricing views when available. 1 source link is attached for reference.

Benchmark Results

GPT-5

Benchmark Results

Thinking
Tool usage

Abstract Generalization

8 evaluations
Benchmark / mode
Score
Rank/total
ARC-AGI-1
Standard Mode
6
171 / 176
ARC-AGI-1
Thinking Level · Low
44
136 / 176
ARC-AGI-1
Thinking Level · Medium
56.17
126 / 176
ARC-AGI-1
Thinking Level · High
65.67
106 / 176
ARC-AGI-2
Standard Mode
0
160 / 164
ARC-AGI-2
Thinking Level · Low
1.94
144 / 164
ARC-AGI-2
Thinking Level · Medium
7.49
117 / 164
ARC-AGI-2
Thinking Level · High
9.86
113 / 164

Knowledge Exams

9 evaluations
Benchmark / mode
Score
Rank/total
HLE
Standard Mode
6.60
456 / 568
HLE
Standard Mode
6.30
461 / 568
HLE
unknown
25.32
251 / 568
HLE
Thinking Level · Low
19.60
296 / 568
HLE
Thinking Level · Medium
25.40
249 / 568
HLE
Thinking Enabled
24.80
257 / 568
HLE
Thinking EnabledTools
35.20
173 / 568
HLE
Thinking Level · High
28.50
224 / 568
HLE
Thinking Level · High
25.32
251 / 568

Scientific Reasoning

7 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
77.80
258 / 461
GPQA Diamond
Thinking Level · Low
80.80
230 / 461
GPQA Diamond
Thinking Level · Medium
85.35
156 / 461
GPQA Diamond
Thinking EnabledTools
87.30
129 / 461
GPQA Diamond
Thinking Level · High
86.17
147 / 461
CritPt
Thinking Level · Low
1.10
151 / 204
CritPt
Thinking Level · High
5.70
92 / 204

Repository Engineering

2 evaluations
Benchmark / mode
Score
Rank/total
SWE-bench Verified
Thinking Level · High
72.80
51 / 116
SWE-Bench Pro - Public
Thinking Level · High
36.30
59 / 62

Algorithmic Coding

7 evaluations
Benchmark / mode
Score
Rank/total
CodeClash
Standard ModeTools
1360
2 / 8
LiveCodeBench
Standard Mode
55.80
150 / 251
LiveCodeBench
Thinking Level · Low
76.30
65 / 251
LiveCodeBench
Thinking Level · Medium
70.30
94 / 251
LiveCodeBench
Thinking Level · High
84.60
33 / 251
IOI 2025 (Vals v1)
Thinking EnabledTools
29
2 / 9
IOI 2024 (Vals v1)
Thinking EnabledTools
11
4 / 10

Mathematics

18 evaluations
Benchmark / mode
Score
Rank/total
AIME2025
Standard Mode
61.90
136 / 215
AIME2025
Thinking Level · Low
83
86 / 215
AIME2025
Thinking Level · Medium
91.70
46 / 215
AIME2025
Thinking Enabled
94.60
33 / 215
AIME2025
Thinking EnabledTools
99.60
10 / 215
AIME2025
Thinking Level · High
94.30
35 / 215
AIME2025
Thinking Level · HighTools
95
30 / 215
IMO-ProofBench
Thinking Enabled
59
2 / 16
FrontierMath v2
Thinking Level · Low
37.19
37 / 58
FrontierMath v2
Thinking Level · High
55.44
27 / 58
FrontierMath
Thinking Level · Medium
24.80
15 / 60
FrontierMath
Thinking Level · High
24.80
15 / 60
FrontierMath
Thinking Level · HighTools
26.30
14 / 60
FrontierMath Tier 4 v2
Thinking Level · High
21.95
26 / 42
IMO-ProofBench Advanced
Thinking Enabled
20
11 / 24
FrontierMath - Tier 4
Thinking Level · Medium
6.30
35 / 80
FrontierMath - Tier 4
Thinking Level · High
12.50
29 / 80
MathArena Apex
Thinking Level · HighTools
1.04
14 / 17

Writing

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1623.50
36 / 106

Agentic Development

9 evaluations
Benchmark / mode
Score
Rank/total
Aider-Polyglot
Thinking Level · Low
81.30
5 / 59
Aider-Polyglot
Thinking Level · Medium
86.70
2 / 59
Aider-Polyglot
Thinking Level · High
88
1 / 59
Terminal-Bench
Thinking EnabledTools
43.80
8 / 35
Terminal Bench Hard
Standard ModeTools
18.20
154 / 244
Terminal Bench Hard
Thinking Level · LowTools
26.50
123 / 244
Terminal Bench Hard
Thinking Level · MediumTools
37.90
59 / 244
Terminal Bench Hard
Thinking Level · HighTools
32.60
93 / 244
Terminal-Bench 2.1
Thinking Level · HighTools
35.20
158 / 199

Visual Understanding

11 evaluations
Benchmark / mode
Score
Rank/total
MMMU
Standard Mode
74.40
34 / 74
MMMU
Thinking Enabled
84.20
6 / 74
MMMU
Thinking Level · High
84.20
6 / 74
GeoBench ACW
Thinking Level · Medium
81
5 / 20
MMMU-Pro
Standard Mode
62.10
167 / 229
MMMU-Pro
Thinking Level · Low
73.80
104 / 229
MMMU-Pro
Thinking Level · Medium
74.30
96 / 229
MMMU-Pro
Thinking Enabled
78.40
61 / 229
MMMU-Pro
Thinking Level · High
74.20
97 / 229
VPCT
Thinking Level · Medium
63.20
6 / 24
VPCT
Thinking Level · High
66
5 / 24

Commonsense

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Thinking Level · High
56.70
43 / 96

Service Workflows

8 evaluations
Benchmark / mode
Score
Rank/total
τ²-Bench - Telecom
Standard ModeTools
67
144 / 264
τ²-Bench - Telecom
Thinking Level · LowTools
84.20
96 / 264
τ²-Bench - Telecom
Thinking Level · MediumTools
86.50
81 / 264
τ²-Bench - Telecom
Thinking EnabledTools
95.83
25 / 264
τ²-Bench - Telecom
Thinking Level · HighTools
84.80
90 / 264
τ²-Bench
Thinking EnabledTools
80
15 / 44
SAGE
Thinking Level · HighTools
43.68
36 / 64
τ³-Banking
Thinking Level · HighTools
22.10
86 / 167

Instruction Following

4 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Standard Mode
45.60
178 / 282
IF Bench
Thinking Level · Low
66.60
92 / 282
IF Bench
Thinking Level · Medium
70.60
67 / 282
IF Bench
Thinking Level · High
73.10
47 / 282

Fact Finding

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
Thinking EnabledTools
54.90
42 / 58

Long Reasoning

3 evaluations
Benchmark / mode
Score
Rank/total
Fiction.liveBench
Thinking Level · Medium
97.20
1 / 16
AA-LCR
Thinking Level · Medium
76
85 / 174
AA-LCR
Thinking Level · High
78.20
69 / 174

Cross-industry Work

1 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA v2
Thinking Level · HighTools
1015
94 / 110

Maintenance & Optimization

1 evaluations
Benchmark / mode
Score
Rank/total
GSO
Thinking Level · HighTools
6.90
16 / 21

ML Engineering

1 evaluations
Benchmark / mode
Score
Rank/total
WeirdML v2
Thinking Level · HighTools
60.70
29 / 52

Preference Arenas

1 evaluations
Benchmark / mode
Score
Rank/total
Text Arena (Coding)
Thinking Level · Medium
1395
29 / 35

Capability Frontier Metrics

2 evaluations
Benchmark / mode
Score
Rank/total
METR Time Horizons v1.1
Thinking Level · MediumTools
137.32
11 / 22
METR Time Horizons v1.1
Thinking Level · HighTools
203.01
9 / 22

Code Generation & Editing

1 evaluations
Benchmark / mode
Score
Rank/total
Vibe Code Bench v1.1
Thinking Level · HighTools
20.09
48 / 60

Clinical Workflows

2 evaluations
Benchmark / mode
Score
Rank/total
MedScribe
Thinking Level · HighTools
83.65
30 / 66
MedCode
Thinking Level · HighTools
49.63
13 / 64

Capability Indices

1 evaluations
Benchmark / mode
Score
Rank/total
ECI
unknown
150
47 / 167

Memory & Persistence

1 evaluations
Benchmark / mode
Score
Rank/total
EBR-bench
Thinking Level · HighTools
12.70
16 / 23

Competitor Comparison

Benchmark scores for GPT-5 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGPT-5CurrentClaude Opus 4Gemini 2.5-Pro
ARC-AGI-1
Score (%)
Abstract Generalization
65.67Thinking Level · High
35.67Standard Mode
37.00Thinking Enabled
ARC-AGI-2
Score (% solved); cost per task (USD)
Abstract Generalization
9.86Thinking Level · High
8.61Standard Mode
4.86Thinking Enabled
HLE
Accuracy
Knowledge Exams
35.20Thinking Enabled | Tools
12.30Thinking Enabled
22.50Thinking Enabled
CritPt
Score
Scientific Reasoning
5.70Thinking Level · High
--
2.60Thinking Enabled
GPQA Diamond
Accuracy
Scientific Reasoning
87.30Thinking Enabled | Tools
79.60Standard Mode
86.40Thinking Enabled
SWE-bench Verified
Accuracy
Repository Engineering
72.80Thinking Level · High
72.50Standard Mode
67.20Thinking Enabled
CodeClash
Elo / win rate
Algorithmic Coding
1360.00Standard Mode | Tools
--
1125.00Standard Mode | Tools
IOI 2024 (Vals v1)
Subtask points (%)
Algorithmic Coding
11.00Thinking Enabled | Tools
--
19.00Thinking Enabled | Tools
IOI 2025 (Vals v1)
Subtask points (%)
Algorithmic Coding
29.00Thinking Enabled | Tools
--
15.20Thinking Enabled | Tools
LiveCodeBench
Pass @K
Algorithmic Coding
84.60Thinking Level · High
63.60Thinking Enabled
80.10Thinking Enabled
AIME2025
Accuracy
Mathematics
99.60Thinking Enabled | Tools
75.50Standard Mode
88.00Thinking Enabled
FrontierMath
Accuracy
Mathematics
26.30Thinking Level · High | Tools
4.50Standard Mode
11.00Standard Mode
22 additional benchmarks remain in the chart above.

Standard API Pricing: GPT-5 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Gemini 2.5-Pro: Base price applies to <= 200000
ModelSupplierStandard inputStandard outputBase price applies to
GPT-5
OpenAI$1.25 / 1M tokens$10 / 1M tokens—
Claude Opus 4
Anthropic$15 / 1M tokens$75 / 1M tokens—
Gemini 2.5-Pro
Google DeepMind$1.25 / 1M tokens$10 / 1M tokens<= 200000

Version History

How each version of the GPT-5 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGPT-5CurrentGPT-4.5GPT-4.1GPT-4o(2025-03-27)
ARC-AGI-1
Score (%)
Abstract Generalization
65.67Thinking Level · High
--
--
8.80Standard Mode
HLE
Accuracy
Knowledge Exams
35.20Thinking Enabled | Tools
5.44Standard Mode
5.40Thinking Level · High
--
GPQA Diamond
Accuracy
Scientific Reasoning
87.30Thinking Enabled | Tools
71.40Standard Mode
66.30Standard Mode
66.90Standard Mode
SWE-bench Verified
Accuracy
Repository Engineering
72.80Thinking Level · High
38.00Standard Mode
54.60Standard Mode
--
LiveCodeBench
Pass @K
Algorithmic Coding
84.60Thinking Level · High
46.40Standard Mode
40.50Standard Mode
35.80Standard Mode
AIME2025
Accuracy
Mathematics
99.60Thinking Enabled | Tools
--
36.70Standard Mode
26.70Standard Mode
FrontierMath
Accuracy
Mathematics
26.30Thinking Level · High | Tools
--
5.50Standard Mode
--
Creative Writing
Elo、大模型评判两两对战
Writing
1623.50Standard Mode
1255.30Standard Mode
1417.20Standard Mode
1502.10Standard Mode
Aider-Polyglot
Accuracy
Agentic Development
88.00Thinking Level · High
44.90Standard Mode
52.40Standard Mode
45.30Standard Mode
Terminal Bench Hard
Accuracy
Agentic Development
37.90Thinking Level · Medium | Tools
--
13.60Standard Mode | Tools
--
GeoBench ACW
ACW country accuracy (%)
Visual Understanding
81.00Thinking Level · Medium
--
72.00Standard Mode
--
MMMU
Accuracy
Visual Understanding
84.20Thinking Level · High
74.40Thinking Level · High
--
--
8 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: ARC-AGI-1 · Abstract Generalization

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the GPT-5 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
GPT-5
OpenAI$1.25 / 1M tokens$10 / 1M tokens—
GPT-4.1
Microsoft Azure$2 / 1M tokens$8 / 1M tokens—
GPT-4o(2025-03-27)
OpenAI$2.5 / 1M tokens$10 / 1M tokens—

Sources