DataLearner logo

GPT-5 Benchmark Details

GPT-5 currently shows benchmark results led by Aider-Polyglot (1 / 59, score 88), Fiction.liveBench (1 / 16, score 97.20), AIME2025 (9 / 106, score 99.60). This page also compares it with 2 competitor models and 3 predecessor or same-series models, including performance and pricing views when available. 1 source link is attached for reference.

Benchmark Results

GPT-5

Benchmark Results

Thinking
Tool usage

General Knowledge

12 evaluations
Benchmark / mode
Score
Rank/total
ARC-AGI-1
Standard Mode
6
87 / 91
ARC-AGI-1
Thinking Level · Low
44
71 / 91
ARC-AGI-1
Thinking Level · Medium
56.20
66 / 91
ARC-AGI-1
Thinking Level · High
65.70
56 / 91
HLE
Standard Mode
6.30
179 / 190
HLE
Thinking Enabled
24.80
119 / 190
HLE
Thinking EnabledTools
35.20
85 / 190
HLE
Thinking Level · High
25.32
117 / 190
ARC-AGI-2
Standard Mode
0
83 / 85
ARC-AGI-2
Thinking Level · Low
1.90
76 / 85
ARC-AGI-2
Thinking Level · Medium
7.50
66 / 85
ARC-AGI-2
Thinking Level · High
9.90
63 / 85

Other

4 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
77.80
157 / 271
GPQA Diamond
Thinking Level · Medium
85.35
97 / 271
GPQA Diamond
Thinking EnabledTools
87.30
80 / 271
GPQA Diamond
Thinking Level · High
86.17
91 / 271

Coding and Software Engineer

6 evaluations
Benchmark / mode
Score
Rank/total
Text Arena (Coding)
Thinking Level · Medium
1395
29 / 35
CodeClash
Standard ModeTools
1360
2 / 8
SWE-bench Verified
Thinking Level · High
72.80
51 / 115
WeirdML v2
Thinking Level · HighTools
60.70
29 / 52
SWE-Bench Pro - Public
Thinking Level · High
36.30
58 / 60
GSO
Thinking Level · HighTools
6.90
16 / 21

Math and Reasoning

15 evaluations
Benchmark / mode
Score
Rank/total
AIME2025
Standard Mode
61.90
80 / 106
AIME2025
Thinking Enabled
94.60
26 / 106
AIME2025
Thinking EnabledTools
99.60
9 / 106
IMO-ProofBench
Thinking Enabled
59
2 / 16
FrontierMath v2
Thinking Level · Low
37.19
37 / 58
FrontierMath v2
Thinking Level · High
55.44
27 / 58
IMO 2025
Thinking Enabled
29
2 / 9
FrontierMath
Thinking Level · Medium
24.80
15 / 60
FrontierMath
Thinking Level · High
24.80
15 / 60
FrontierMath
Thinking Level · HighTools
26.30
14 / 60
FrontierMath Tier 4 v2
Thinking Level · High
21.95
25 / 41
IMO-ProofBench Advanced
Thinking Enabled
20
9 / 19
FrontierMath - Tier 4
Thinking Level · Medium
6.30
35 / 80
FrontierMath - Tier 4
Thinking Level · High
12.50
29 / 80
IMO 2024
Thinking Enabled
11
4 / 10

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1623.40
32 / 99

AI Agent - Tool Usage

1 evaluations
Benchmark / mode
Score
Rank/total
Terminal-Bench
Thinking EnabledTools
43.80
8 / 35

Multimodal Understanding

4 evaluations
Benchmark / mode
Score
Rank/total
MMMU
Thinking Level · High
84.20
5 / 28
GeoBench ACW
Thinking Level · Medium
81
5 / 20
VPCT
Thinking Level · Medium
63.20
6 / 24
VPCT
Thinking Level · High
66
5 / 24

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Thinking Level · High
56.70
38 / 92

Agent Level Benchmark

8 evaluations
Benchmark / mode
Score
Rank/total
METR Time Horizons v1.1
Thinking Level · MediumTools
137.32
11 / 22
METR Time Horizons v1.1
Thinking Level · HighTools
203.01
9 / 22
τ²-Bench - Telecom
Thinking EnabledTools
95.80
13 / 35
τ²-Bench - Telecom
Thinking Level · HighTools
96.70
11 / 35
Aider-Polyglot
Thinking Level · Low
81.30
5 / 59
Aider-Polyglot
Thinking Level · Medium
86.70
2 / 59
Aider-Polyglot
Thinking Level · High
88
1 / 59
τ²-Bench
Thinking EnabledTools
80
15 / 44

Instruction Following

1 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Thinking Level · High
73.10
13 / 35

AI Agent - Information Search

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
Thinking EnabledTools
54.90
40 / 56

Other

1 evaluations
Benchmark / mode
Score
Rank/total
Fiction.liveBench
Thinking Level · Medium
97.20
1 / 16

Competitor Comparison

Benchmark scores for GPT-5 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGPT-5CurrentClaude Opus 4Gemini 2.5-Pro
ARC-AGI-1
综合评估
65.70Thinking Level · High
35.70Standard Mode
37.00Thinking Enabled
ARC-AGI-2
综合评估
9.90Thinking Level · High
8.60Standard Mode
4.90Thinking Enabled
HLE
综合评估
35.20Thinking Enabled | Tools
10.70Standard Mode
21.60Thinking Enabled
GPQA Diamond
科学与综合推理
87.30Thinking Enabled | Tools
79.60Standard Mode
86.40Thinking Enabled
CodeClash
编程与软件工程
1360.00Standard Mode | Tools
--
1125.00Standard Mode | Tools
GSO
编程与软件工程
6.90Thinking Level · High | Tools
--
3.92Standard Mode | Tools
SWE-bench Verified
编程与软件工程
72.80Thinking Level · High
72.50Standard Mode
67.20Thinking Enabled
AIME2025
数学推理
99.60Thinking Enabled | Tools
75.50Standard Mode
88.00Thinking Enabled
FrontierMath
数学推理
26.30Thinking Level · High | Tools
4.50Standard Mode
11.00Standard Mode
12.50Thinking Level · High
4.2032K
2.10Standard Mode
IMO 2024
数学推理
11.00Thinking Enabled
--
19.00Thinking Enabled
IMO 2025
数学推理
29.00Thinking Enabled
--
15.20Thinking Enabled
14 additional benchmarks remain in the chart above.

Standard API Pricing: GPT-5 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Gemini 2.5-Pro: Base price applies to <= 200000
ModelSupplierStandard inputStandard outputBase price applies to
GPT-5
OpenAI$1.25 / 1M tokens$10 / 1M tokens
Claude Opus 4
Anthropic$15 / 1M tokens$75 / 1M tokens
Gemini 2.5-Pro
Google Deep Mind$1.25 / 1M tokens$10 / 1M tokens<= 200000

Version History

How each version of the GPT-5 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGPT-5CurrentGPT-4.5GPT-4.1GPT-4o(2025-03-27)
ARC-AGI-1
综合评估
65.70Thinking Level · High
--
--
8.80Standard Mode
HLE
综合评估
35.20Thinking Enabled | Tools
--
3.70Standard Mode
--
GPQA Diamond
科学与综合推理
87.30Thinking Enabled | Tools
71.40Standard Mode
66.30Standard Mode
66.90Standard Mode
SWE-bench Verified
编程与软件工程
72.80Thinking Level · High
38.00Standard Mode
54.60Standard Mode
--
WeirdML v2
编程与软件工程
60.70Thinking Level · High | Tools
--
39.04Standard Mode | Tools
--
AIME2025
数学推理
99.60Thinking Enabled | Tools
--
36.70Standard Mode
26.70Standard Mode
FrontierMath
数学推理
26.30Thinking Level · High | Tools
--
5.50Standard Mode
--
Creative Writing
写作和创作
1623.40Standard Mode
1255.00Standard Mode
1417.00Standard Mode
1502.00Standard Mode
GeoBench ACW
多模态理解
81.00Thinking Level · Medium
--
72.00Standard Mode
--
SimpleBench
常识推理
56.70Thinking Level · High
34.50Standard Mode
27.00Standard Mode
--
Aider-Polyglot
Agent能力评测
88.00Thinking Level · High
44.90Standard Mode
52.40Standard Mode
45.30Standard Mode
τ²-Bench
Agent能力评测
80.00Thinking Enabled | Tools
--
54.70Standard Mode | Tools
--
1 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: ARC-AGI-1 · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the GPT-5 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
GPT-5
OpenAI$1.25 / 1M tokens$10 / 1M tokens
GPT-4.1
Microsoft Azure$2 / 1M tokens$8 / 1M tokens
GPT-4o(2025-03-27)
OpenAI$2.5 / 1M tokens$10 / 1M tokens

Sources