DataLearner logo

GPT-5.6 Sol Benchmark Details

GPT-5.6 Sol currently shows benchmark results led by TerminalBench 2.1 (1 / 29, score 88.80), DeepSWE (1 / 20, score 72.70), SWE-Bench Pro - Public (6 / 55, score 64.60). This page also compares it with 2 competitor models and 4 predecessor or same-series models, including performance and pricing views when available. 3 source links are attached for reference.

Benchmark Results

GPT-5.6 Sol

Benchmark Results

Thinking
Tool usage

Coding and Software Engineer

3 evaluations
Benchmark / mode
Score
Rank/total
AA Coding Agent Index
Extra-HighTools
80
1 / 3
DeepSWE
Extra-HighTools
72.70
1 / 20
SWE-Bench Pro - Public
Extra-HighTools
64.60
6 / 55

AI Agent - Tool Usage

3 evaluations
Benchmark / mode
Score
Rank/total
88.80
1 / 29
Vals CyberBench
Extra-HighTools
88.14
1 / 1
OSWorld 2.0
Extra-HighTools
62.60
2 / 3

General Knowledge

2 evaluations
Benchmark / mode
Score
Rank/total
Vals Index
Extra-High
72.63
1 / 2

Agent Level Benchmark

1 evaluations
Benchmark / mode
Score
Rank/total
Agents' Last Exam
Extra-HighTools
52.70
1 / 5

Competitor Comparison

Benchmark scores for GPT-5.6 Sol compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

3 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGPT-5.6 SolCurrentClaude Fable 5Claude Opus 4.8
DeepSWE
编程与软件工程
72.70Thinking Level · Extra High | Tools
70.00Deep Thinking Mode | Tools
59.00Deep Thinking Mode | Tools
SWE-Bench Pro - Public
编程与软件工程
64.60Thinking Level · Extra High | Tools
80.30Deep Thinking Mode | Tools
69.20Extended Thinking | Tools
TerminalBench 2.1
AI Agent - 工具使用
88.80Thinking Level · High
88.00Thinking Level · High | Tools
78.90Thinking Level · High | Tools

Standard API Pricing: GPT-5.6 Sol vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
GPT-5.6 Sol
OpenAI$5 / 1M tokens$30 / 1M tokens
Claude Fable 5
Anthropic$10 / 1M tokens$50 / 1M tokens
Claude Opus 4.8
Anthropic$5 / 1M tokens$25 / 1M tokens

Version History

How each version of the GPT-5.6 Sol series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

3 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGPT-5.6 SolCurrentGPT-5.5GPT-5.4GPT-5.2
DeepSWE
编程与软件工程
72.70Thinking Level · Extra High | Tools
67.00Thinking Level · Extra High | Tools
52.00Thinking Level · Extra High | Tools
--
SWE-Bench Pro - Public
编程与软件工程
64.60Thinking Level · Extra High | Tools
58.60Thinking Level · High | Tools
57.70Thinking Level · Extra High
55.60Thinking Level · Extra High | Tools
TerminalBench 2.1
AI Agent - 工具使用
88.80Thinking Level · High
83.40Thinking Level · High | Tools
--
--

Single-Benchmark Version Trend

Viewing: DeepSWE · 编程与软件工程

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the GPT-5.6 Sol Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

GPT-5.4: Base price applies to <= 272K
ModelSupplierStandard inputStandard outputBase price applies to
GPT-5.6 Sol
OpenAI$5 / 1M tokens$30 / 1M tokens
GPT-5.5
OpenAI$5 / 1M tokens$30 / 1M tokens
GPT-5.4
OpenAI$2.5 / 1M tokens$15 / 1M tokens<= 272K
GPT-5.2
Facebook AI研究实验室$1.75 / 1M tokens$14 / 1M tokens

Sources