GPT-5.6 Sol Benchmark Details
GPT-5.6 Sol currently shows benchmark results led by Terminal Bench Hard (1 / 244, score 65.90), CritPt (1 / 204, score 32.30), GPQA Diamond (2 / 253, score 94.60). This page also compares it with 4 competitor models and 4 predecessor or same-series models, including performance and pricing views when available. 4 source links are attached for reference.
Benchmark Results
Benchmark Results
Abstract Generalization
15 evaluationsScientific Reasoning
10 evaluationsRepository Engineering
10 evaluationsAgentic Development
18 evaluationsMemory & Persistence
2 evaluationsTool Orchestration
7 evaluationsCapability Indices
8 evaluationsCoding Indices
2 evaluationsVulnerability Analysis
4 evaluationsDesktop Workflows
2 evaluationsCode Generation & Editing
4 evaluationsOffice & Business
2 evaluationsVisual Understanding
8 evaluationsScientific Computing
6 evaluationsService Workflows
8 evaluationsLegal
3 evaluationsFinance
3 evaluationsMathematics
5 evaluationsExploitation
5 evaluationsDocuments & Charts
6 evaluationsData Analysis
2 evaluationsDesign & Engineering
2 evaluationsMaintenance & Optimization
4 evaluationsBiology & Genomics
2 evaluationsChemistry & Drug Discovery
1 evaluationsMedical Reasoning
1 evaluationsSafe Computer Use
2 evaluationsPolicy Compliance
2 evaluationsLong Retrieval
2 evaluationsClinical Workflows
2 evaluationsCompetitor Comparison
Benchmark scores for GPT-5.6 Sol compared against top models in its class
12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.
| Benchmark | GPT-5.6 SolCurrent | Claude Fable 5 | Kimi K3 | Claude Opus 4.8 | GLM-5.2 |
|---|---|---|---|---|---|
97.50Thinking Level · Extra High | 98.50Thinking Level · Extra High | Tools | 94.50Thinking Level · High | Tools | 92.50Thinking Level · High | Tools | -- | |
92.50Thinking Level · High | 89.17Thinking Level · High | 60.42Thinking Level · High | Tools | 72.08Thinking Level · High | Tools | -- | |
ARC-AGI-3 (Standard harness) Action efficiency score(以 ARC Prize Standard harness 口径为准) Abstract Generalization | 7.78Thinking Level · High | -- | -- | 1.52Thinking Level · High | -- |
44.50Thinking Level · High | 59.00Deep Thinking Mode | 59.80Thinking Level · High | Tools | 57.90Extended Thinking | Tools | 54.70Thinking Enabled | Tools | |
32.30Thinking Level · High | 28.60Thinking Level · High | 23.40Thinking Level · High | 20.90Thinking Level · High | 20.90Thinking Level · High | |
94.60Thinking Level · High | 85.86Thinking Level · High | 91.92Thinking Level · High | 93.60Thinking Level · High | 91.86Thinking Level · High | |
1963.40Standard Mode | 1934.60Standard Mode | 2070.60Standard Mode | 1835.20Standard Mode | 1752.80Standard Mode | |
64.80Thinking Level · Extra High | 81.90Standard Mode | 60.70Thinking Level · High | 64.80Standard Mode | 58.80Standard Mode | |
72.70Thinking Level · Extra High | Tools | 69.91Deep Thinking Mode | Tools | 68.51Thinking Level · High | Tools | 58.97Deep Thinking Mode | Tools | 44.00Deep Thinking Mode | Tools | |
56.80Thinking Level · High | Tools | -- | 58.00Thinking Level · High | Tools | -- | 48.90Thinking Enabled | Tools | |
64.60Thinking Level · Extra High | Tools | 80.30Deep Thinking Mode | Tools | -- | 69.20Extended Thinking | Tools | 62.10Thinking Enabled | Tools | |
95.50Thinking Level · Extra High | Tools | -- | 97.70Thinking Level · High | Tools | -- | -- |
Standard API Pricing: GPT-5.6 Sol vs. Peer Models
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier.
These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
GPT-5.6 Sol | OpenAI | $4 / 1M tokens | $20 / 1M tokens | — |
Claude Fable 5 | Anthropic | $10 / 1M tokens | $50 / 1M tokens | — |
Kimi K3 | Moonshot AI | ¥20 / 1M tokens | ¥100 / 1M tokens | — |
Claude Opus 4.8 | Anthropic | $5 / 1M tokens | $25 / 1M tokens | — |
GLM-5.2 | 智谱AI | $1.4 / 1M tokens | $4.4 / 1M tokens | — |
Version History
How each version of the GPT-5.6 Sol series stacks up on benchmark tests
12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.
| Benchmark | GPT-5.6 SolCurrent | GPT-5.5 | GPT-5.4 | GPT-5.2 |
|---|---|---|---|---|
97.50Thinking Level · Extra High | 95.00Thinking Level · Extra High | 93.67Thinking Level · Extra High | 90.50Deep Thinking Mode | |
92.50Thinking Level · High | 85.00Thinking Level · Extra High | 77.10Standard Mode | 54.20Deep Thinking Mode | |
ARC-AGI-3 (Standard harness) Action efficiency score(以 ARC Prize Standard harness 口径为准) Abstract Generalization | 7.78Thinking Level · High | 0.43Thinking Level · High | 0.21Thinking Level · High | -- |
44.50Thinking Level · High | 52.20Thinking Level · High | Tools | 52.10Thinking Level · Extra High | Tools | 45.50Deep Thinking Mode | Tools | |
32.30Thinking Level · High | 27.10Thinking Level · Extra High | 23.40Thinking Level · Extra High | 11.60Thinking Level · Extra High | |
94.60Thinking Level · High | 94.00Thinking Level · Extra High | 89.90Thinking Level · High | 93.20Deep Thinking Mode | |
1963.40Standard Mode | 1843.50Standard Mode | 1835.60Standard Mode | 1699.80Standard Mode | |
64.80Thinking Level · Extra High | 69.00Standard Mode | -- | 45.80Thinking Level · High | |
72.70Thinking Level · Extra High | Tools | 67.04Thinking Level · Extra High | Tools | 51.77Thinking Level · Extra High | Tools | -- | |
64.60Thinking Level · Extra High | Tools | 58.60Thinking Level · High | Tools | 57.70Thinking Level · Extra High | 55.60Thinking Level · Extra High | Tools | |
90.40Thinking Level · High | 84.40Thinking Level · High | Tools | 82.70Thinking Level · Extra High | Tools | 65.80Thinking Level · Extra High | Tools | |
65.90Thinking Level · High | Tools | 60.60Thinking Level · Extra High | Tools | 57.60Thinking Level · Extra High | Tools | 47.00Thinking Level · Extra High | Tools |
Single-Benchmark Version Trend
Viewing: ARC-AGI-1 · Abstract Generalization
Standard API Pricing Across the GPT-5.6 Sol Series
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
GPT-5.6 Sol | OpenAI | $4 / 1M tokens | $20 / 1M tokens | — |
GPT-5.5 | OpenAI | $5 / 1M tokens | $30 / 1M tokens | — |
GPT-5.4 | OpenAI | $2.5 / 1M tokens | $15 / 1M tokens | — |
GPT-5.2 | OpenAI | $1.75 / 1M tokens | $14 / 1M tokens | — |


