DataLearner logo

GPT-5.6 Sol Benchmark Details

GPT-5.6 Sol currently shows benchmark results led by Terminal Bench Hard (1 / 244, score 65.90), CritPt (1 / 204, score 32.30), GPQA Diamond (2 / 253, score 94.60). This page also compares it with 4 competitor models and 4 predecessor or same-series models, including performance and pricing views when available. 4 source links are attached for reference.

Benchmark Results

GPT-5.6 Sol

Benchmark Results

Thinking
Tool usage
Internet

Abstract Generalization

15 evaluations
Benchmark / mode
Score
Rank/total
74.50
98 / 176
ARC-AGI-1
Medium
92.50
45 / 176
97
21 / 176
96.50
22 / 176
ARC-AGI-1
Extra-High
97.50
9 / 176
42.50
86 / 164
ARC-AGI-2
Medium
67.08
54 / 164
85.42
22 / 164
92.50
4 / 164
ARC-AGI-2
Extra-High
90
10 / 164
6.99
10 / 42

Knowledge Exams

2 evaluations
Benchmark / mode
Score
Rank/total
54.50
2 / 6
HLE
Max
44.50
61 / 233

Scientific Reasoning

10 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
82.83
109 / 253
94.60
2 / 253
89.90
44 / 253
33.33
7 / 13
CritPt
Standard Mode
5.10
98 / 204
14.90
60 / 204
CritPt
Medium
22.90
33 / 204
CritPt
High
25.70
26 / 204
32.30
1 / 204
CritPt
Extra-High
28.60
16 / 204

Writing

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1963.40
6 / 106

Commonsense

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Extra-High
64.80
24 / 96

Repository Engineering

10 evaluations
Benchmark / mode
Score
Rank/total
SWE-Bench Pro V2
Extra-HighTools
95.50
6 / 10
SWE-Bench Pro V2 Hard
Extra-HighTools
82.40
8 / 11
DeepSWE
LowTools
45.35
71 / 91
DeepSWE
MediumTools
61.06
47 / 91
DeepSWE
HighTools
69.40
21 / 91
DeepSWE
MaxTools
72.67
13 / 91
DeepSWE
Extra-HighTools
72.70
12 / 91
SWE-Bench Pro - Public
Extra-HighTools
64.60
9 / 62
56.80
9 / 16
53.50
3 / 6

Fact Finding

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
unknown
90.40
4 / 58

Agentic Development

18 evaluations
Benchmark / mode
Score
Rank/total
88.80
5 / 57
67.20
4 / 5
60.60
7 / 244
62.90
2 / 244
62.10
5 / 244
65.90
1 / 244
Terminal Bench Hard
Extra-HighTools
61.40
6 / 244
24.60
42 / 46
CursorBench 4.0
MediumTools
31.10
33 / 46
35.70
25 / 46
41.70
13 / 46
CursorBench 4.0
Extra-HighTools
37.70
20 / 46
1
83 / 95
14.60
43 / 95
20.70
34 / 95
37.27
20 / 95
Terminal-Bench 4.0
Extra-HighTools
24.70
31 / 95
34.40
2 / 11

Memory & Persistence

2 evaluations
Benchmark / mode
Score
Rank/total
97.63
2 / 126
EBR-bench
MaxTools
44.76
6 / 23

Tool Orchestration

7 evaluations
Benchmark / mode
Score
Rank/total
92.91
2 / 8
84.19
7 / 45
MCP-Atlas
MaxTools
81.80
13 / 44
APEX-Agents
MaxTools
56.70
3 / 7
53.60
3 / 24
26.70
17 / 24
Agents' Last Exam
Extra-HighTools
52.70
4 / 24

Capability Indices

8 evaluations
Benchmark / mode
Score
Rank/total

Coding Indices

2 evaluations
Benchmark / mode
Score
Rank/total

Vulnerability Analysis

4 evaluations
Benchmark / mode
Score
Rank/total
Vals CyberBench
Extra-HighTools
88.14
1 / 1
CyberGym
MaxTools
84.50
5 / 11
79.10
2 / 6
74.30
3 / 6

Desktop Workflows

2 evaluations
Benchmark / mode
Score
Rank/total
65.70
6 / 13
OSWorld 2.0
Extra-HighTools
62.60
8 / 13

Code Generation & Editing

4 evaluations
Benchmark / mode
Score
Rank/total
80.50
16 / 60
60.60
5 / 7
47.50
8 / 9
23
8 / 15

Office & Business

2 evaluations
Benchmark / mode
Score
Rank/total
18.10
23 / 23
45.80
10 / 23

Visual Understanding

8 evaluations
Benchmark / mode
Score
Rank/total
BabyVision
MaxTools
88.90
4 / 9
MMMU-Pro
Standard Mode
71.90
118 / 229
81
35 / 229
MMMU-Pro
Medium
81.40
31 / 229
81.80
27 / 229
83.40
16 / 229
MMMU-Pro
Extra-High
82.70
19 / 229
53
1 / 6

Scientific Computing

6 evaluations
Benchmark / mode
Score
Rank/total
56.40
27 / 134
SciCode
Medium
57.40
18 / 134
57.80
16 / 134
57.10
21 / 134
SciCode
Extra-High
57.10
21 / 134
22.40
5 / 12

Service Workflows

8 evaluations
Benchmark / mode
Score
Rank/total
66.51
10 / 26
SAGE
MaxTools
52.56
4 / 64
τ³-Banking
Standard ModeTools
19.60
95 / 167
τ³-Banking
LowTools
29.10
65 / 167
τ³-Banking
MediumTools
36.50
45 / 167
τ³-Banking
HighTools
36.70
43 / 167
τ³-Banking
MaxTools
44.30
20 / 167
τ³-Banking
Extra-HighTools
46.91
13 / 167

Legal

3 evaluations
Benchmark / mode
Score
Rank/total
87.18
17 / 44
48.08
7 / 42

Finance

3 evaluations
Benchmark / mode
Score
Rank/total
72.34
5 / 42
67.95
8 / 20
53.76
20 / 42

Mathematics

5 evaluations
Benchmark / mode
Score
Rank/total
89.12
1 / 58
83
3 / 42
83
7 / 28
0
2 / 5

ML Engineering

2 evaluations
Benchmark / mode
Score
Rank/total
WeirdML v2
HighTools
88.76
2 / 52
WeirdML v3
Extra-HighTools
0.15
7 / 13

Preference Arenas

1 evaluations
Benchmark / mode
Score
Rank/total
1620.27
4 / 35

Exploitation

5 evaluations
Benchmark / mode
Score
Rank/total

Documents & Charts

6 evaluations
Benchmark / mode
Score
Rank/total
Chartography
MaxTools
79.90
4 / 9
21
33 / 122
GDP.pdf
Medium
26.20
14 / 122
27.80
8 / 122
27.20
10 / 122
GDP.pdf
Extra-High
27.60
9 / 122

Data Analysis

2 evaluations
Benchmark / mode
Score
Rank/total
AA-AnalystAgent
MaxToolsInternet
47.50
7 / 29

GUI Grounding

1 evaluations
Benchmark / mode
Score
Rank/total
76.90
2 / 2

Design & Engineering

2 evaluations
Benchmark / mode
Score
Rank/total
BenchCAD
unknown
83.30
2 / 2
47.40
2 / 2

Music Creation

1 evaluations
Benchmark / mode
Score
Rank/total

Maintenance & Optimization

4 evaluations
Benchmark / mode
Score
Rank/total
SRE-Bench
unknown
68.70
3 / 5
SRE-Bench
unknown
55.90
4 / 5
52.92
7 / 43

Biology & Genomics

2 evaluations
Benchmark / mode
Score
Rank/total
59.90
2 / 2
32.30
2 / 2

Chemistry & Drug Discovery

1 evaluations
Benchmark / mode
Score
Rank/total
47.40
2 / 2

Medical Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total

Safe Computer Use

2 evaluations
Benchmark / mode
Score
Rank/total

Policy Compliance

2 evaluations
Benchmark / mode
Score
Rank/total
48.20
2 / 2
0.29
2 / 2

Honesty & Factuality

1 evaluations
Benchmark / mode
Score
Rank/total
12.20
2 / 2

Long Retrieval

2 evaluations
Benchmark / mode
Score
Rank/total

Algorithmic Coding

1 evaluations
Benchmark / mode
Score
Rank/total
91.17
3 / 26

Clinical Workflows

2 evaluations
Benchmark / mode
Score
Rank/total
MedScribe
MaxTools
85.23
20 / 66
MedCode
MaxTools
43.97
32 / 64

Competitor Comparison

Benchmark scores for GPT-5.6 Sol compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGPT-5.6 SolCurrentClaude Fable 5Kimi K3Claude Opus 4.8GLM-5.2
ARC-AGI-1
Score (%)
Abstract Generalization
97.50Thinking Level · Extra High
98.50Thinking Level · Extra High | Tools
94.50Thinking Level · High | Tools
92.50Thinking Level · High | Tools
--
ARC-AGI-2
Score (% solved); cost per task (USD)
Abstract Generalization
92.50Thinking Level · High
89.17Thinking Level · High
60.42Thinking Level · High | Tools
72.08Thinking Level · High | Tools
--
ARC-AGI-3 (Standard harness)
Action efficiency score(以 ARC Prize Standard harness 口径为准)
Abstract Generalization
7.78Thinking Level · High
--
--
1.52Thinking Level · High
--
HLE
Accuracy
Knowledge Exams
44.50Thinking Level · High
59.00Deep Thinking Mode
59.80Thinking Level · High | Tools
57.90Extended Thinking | Tools
54.70Thinking Enabled | Tools
CritPt
Score
Scientific Reasoning
32.30Thinking Level · High
28.60Thinking Level · High
23.40Thinking Level · High
20.90Thinking Level · High
20.90Thinking Level · High
GPQA Diamond
Accuracy
Scientific Reasoning
94.60Thinking Level · High
85.86Thinking Level · High
91.92Thinking Level · High
93.60Thinking Level · High
91.86Thinking Level · High
Creative Writing
Elo、大模型评判两两对战
Writing
1963.40Standard Mode
1934.60Standard Mode
2070.60Standard Mode
1835.20Standard Mode
1752.80Standard Mode
SimpleBench
Score (AVG@5)
Commonsense
64.80Thinking Level · Extra High
81.90Standard Mode
60.70Thinking Level · High
64.80Standard Mode
58.80Standard Mode
DeepSWE
Pass@1 (DeepSWE v1.1)
Repository Engineering
72.70Thinking Level · Extra High | Tools
69.91Deep Thinking Mode | Tools
68.51Thinking Level · High | Tools
58.97Deep Thinking Mode | Tools
44.00Deep Thinking Mode | Tools
NL2Repo-Bench
Average test pass rate
Repository Engineering
56.80Thinking Level · High | Tools
--
58.00Thinking Level · High | Tools
--
48.90Thinking Enabled | Tools
SWE-Bench Pro - Public
Accuracy
Repository Engineering
64.60Thinking Level · Extra High | Tools
80.30Deep Thinking Mode | Tools
--
69.20Extended Thinking | Tools
62.10Thinking Enabled | Tools
SWE-Bench Pro V2
Resolve Rate (%, pass@1)
Repository Engineering
95.50Thinking Level · Extra High | Tools
--
97.70Thinking Level · High | Tools
--
--
52 additional benchmarks remain in the chart above.

Standard API Pricing: GPT-5.6 Sol vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

GPT-5.6 Sol
Supplier: OpenAI
Standard input: $4 / 1M tokens
Standard output: $20 / 1M tokens
Claude Fable 5
Supplier: Anthropic
Standard input: $10 / 1M tokens
Standard output: $50 / 1M tokens
Kimi K3
Supplier: Moonshot AI
Standard input: ¥20 / 1M tokens
Standard output: ¥100 / 1M tokens
Claude Opus 4.8
Supplier: Anthropic
Standard input: $5 / 1M tokens
Standard output: $25 / 1M tokens
GLM-5.2
Supplier: 智谱AI
Standard input: $1.4 / 1M tokens
Standard output: $4.4 / 1M tokens
ModelSupplierStandard inputStandard outputBase price applies to
GPT-5.6 Sol
OpenAI$4 / 1M tokens$20 / 1M tokens—
Claude Fable 5
Anthropic$10 / 1M tokens$50 / 1M tokens—
Kimi K3
Moonshot AI¥20 / 1M tokens¥100 / 1M tokens—
Claude Opus 4.8
Anthropic$5 / 1M tokens$25 / 1M tokens—
GLM-5.2
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens—

Version History

How each version of the GPT-5.6 Sol series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGPT-5.6 SolCurrentGPT-5.5GPT-5.4GPT-5.2
ARC-AGI-1
Score (%)
Abstract Generalization
97.50Thinking Level · Extra High
95.00Thinking Level · Extra High
93.67Thinking Level · Extra High
90.50Deep Thinking Mode
ARC-AGI-2
Score (% solved); cost per task (USD)
Abstract Generalization
92.50Thinking Level · High
85.00Thinking Level · Extra High
77.10Standard Mode
54.20Deep Thinking Mode
ARC-AGI-3 (Standard harness)
Action efficiency score(以 ARC Prize Standard harness 口径为准)
Abstract Generalization
7.78Thinking Level · High
0.43Thinking Level · High
0.21Thinking Level · High
--
HLE
Accuracy
Knowledge Exams
44.50Thinking Level · High
52.20Thinking Level · High | Tools
52.10Thinking Level · Extra High | Tools
45.50Deep Thinking Mode | Tools
CritPt
Score
Scientific Reasoning
32.30Thinking Level · High
27.10Thinking Level · Extra High
23.40Thinking Level · Extra High
11.60Thinking Level · Extra High
GPQA Diamond
Accuracy
Scientific Reasoning
94.60Thinking Level · High
94.00Thinking Level · Extra High
89.90Thinking Level · High
93.20Deep Thinking Mode
Creative Writing
Elo、大模型评判两两对战
Writing
1963.40Standard Mode
1843.50Standard Mode
1835.60Standard Mode
1699.80Standard Mode
SimpleBench
Score (AVG@5)
Commonsense
64.80Thinking Level · Extra High
69.00Standard Mode
--
45.80Thinking Level · High
DeepSWE
Pass@1 (DeepSWE v1.1)
Repository Engineering
72.70Thinking Level · Extra High | Tools
67.04Thinking Level · Extra High | Tools
51.77Thinking Level · Extra High | Tools
--
SWE-Bench Pro - Public
Accuracy
Repository Engineering
64.60Thinking Level · Extra High | Tools
58.60Thinking Level · High | Tools
57.70Thinking Level · Extra High
55.60Thinking Level · Extra High | Tools
BrowseComp
Accuracy
Fact Finding
90.40Thinking Level · High
84.40Thinking Level · High | Tools
82.70Thinking Level · Extra High | Tools
65.80Thinking Level · Extra High | Tools
Terminal Bench Hard
Accuracy
Agentic Development
65.90Thinking Level · High | Tools
60.60Thinking Level · Extra High | Tools
57.60Thinking Level · Extra High | Tools
47.00Thinking Level · Extra High | Tools
30 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: ARC-AGI-1 · Abstract Generalization

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the GPT-5.6 Sol Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
GPT-5.6 Sol
OpenAI$4 / 1M tokens$20 / 1M tokens—
GPT-5.5
OpenAI$5 / 1M tokens$30 / 1M tokens—
GPT-5.4
OpenAI$2.5 / 1M tokens$15 / 1M tokens—
GPT-5.2
OpenAI$1.75 / 1M tokens$14 / 1M tokens—

Sources