DataLearner logo

GPT-6 Sol Benchmark Details

GPT-6 Sol currently shows benchmark results led by CritPt (6 / 204, score 30.86), Agents' Last Exam (2 / 24, score 56.40), Code Migration (4 / 43, score 57.20). This page also compares it with 2 competitor models and 3 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

GPT-6 Sol

Benchmark Results

Thinking
Tool usage

Commonsense

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Standard Mode
73.10
16 / 96

Agentic Development

2 evaluations
Benchmark / mode
Score
Rank/total
83.15
24 / 57
43.94
17 / 96

Repository Engineering

1 evaluations
Benchmark / mode
Score
Rank/total
DeepSWE
MaxTools
68.80
25 / 91

Capability Indices

1 evaluations
Benchmark / mode
Score
Rank/total
Vals Index
MaxTools
62.57
9 / 42

Desktop Workflows

1 evaluations
Benchmark / mode
Score
Rank/total
OSWorld 2.0
MaxTools
64.40
7 / 13

Tool Orchestration

1 evaluations
Benchmark / mode
Score
Rank/total
56.40
2 / 24

Code Generation & Editing

3 evaluations
Benchmark / mode
Score
Rank/total
87.82
6 / 60
49.30
7 / 9
2
14 / 15

Office & Business

1 evaluations
Benchmark / mode
Score
Rank/total
AutomationBench
Extra-HighTools
33.20
15 / 23

Scientific Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
30.86
6 / 204

Scientific Computing

1 evaluations
Benchmark / mode
Score
Rank/total
57.64
17 / 134

Finance

3 evaluations
Benchmark / mode
Score
Rank/total
71.53
8 / 42
53.05
20 / 20
49.05
31 / 42

Legal

2 evaluations
Benchmark / mode
Score
Rank/total
28.85
32 / 42

Documents & Charts

1 evaluations
Benchmark / mode
Score
Rank/total
24.80
18 / 122

Biology & Genomics

1 evaluations
Benchmark / mode
Score
Rank/total
74.81
2 / 4

Medical Reasoning

4 evaluations
Benchmark / mode
Score
Rank/total

ML Engineering

1 evaluations
Benchmark / mode
Score
Rank/total
WeirdML v3
Extra-HighTools
0.20
4 / 13

Maintenance & Optimization

1 evaluations
Benchmark / mode
Score
Rank/total
57.20
4 / 43

Algorithmic Coding

1 evaluations
Benchmark / mode
Score
Rank/total
82.61
7 / 26

Mathematics

1 evaluations
Benchmark / mode
Score
Rank/total
83
7 / 28

Service Workflows

2 evaluations
Benchmark / mode
Score
Rank/total
56.63
22 / 26
SAGE
MaxTools
44.79
32 / 64

Clinical Workflows

2 evaluations
Benchmark / mode
Score
Rank/total
MedScribe
MaxTools
82.03
37 / 66
MedCode
MaxTools
47.07
27 / 64

Memory & Persistence

1 evaluations
Benchmark / mode
Score
Rank/total
EBR-bench
MaxTools
53.33
4 / 23

Competitor Comparison

Benchmark scores for GPT-6 Sol compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGPT-6 SolCurrentGPT-6 AstraClaude Opus 5.5
SimpleBench
Score (AVG@5)
Commonsense
73.10Standard Mode
83.60Thinking Level · High
88.40Standard Mode
Terminal-Bench 2.1
Accuracy
Agentic Development
83.15Thinking Level · High | Tools
87.42Thinking Level · High | Tools
87.64Thinking Level · High | Tools
Terminal-Bench 4.0
Resolution rate (%)
Agentic Development
43.94Thinking Level · High | Tools
58.18Thinking Level · High | Tools
66.40Thinking Level · Extra High | Tools
DeepSWE
Pass@1 (DeepSWE v1.1)
Repository Engineering
68.80Thinking Level · High | Tools
74.12Thinking Level · Extra High | Tools
--
Vals Index
跨行业任务准确率综合指数
Capability Indices
62.57Standard Mode | Tools
66.61Thinking Level · High | Tools
69.69Standard Mode | Tools
OSWorld 2.0
Partial score
Desktop Workflows
64.40Thinking Level · High | Tools
72.60Thinking Level · High | Tools
81.80Thinking Level · High | Tools
Agents' Last Exam
Score
Tool Orchestration
56.40Thinking Level · High | Tools
59.30Thinking Level · High | Tools
--
FrontierCode 1.1 Main
Score (%)
Code Generation & Editing
49.30Thinking Level · High | Tools
53.30Thinking Level · High | Tools
54.60Thinking Level · Medium | Tools
Program Bench
Score
Code Generation & Editing
2.00Thinking Level · High | Tools
--
18.50Thinking Level · High | Tools
Vibe Code Bench v1.1
Pass rate (%)
Code Generation & Editing
87.82Thinking Level · High | Tools
89.59Thinking Level · High | Tools
90.29Thinking Level · High | Tools
AutomationBench
Pass Rate
Office & Business
33.20Thinking Level · Extra High | Tools
41.40Thinking Level · High | Tools
40.00Thinking Level · High | Tools
CritPt
Score
Scientific Reasoning
30.86Thinking Level · High
31.70Thinking Level · High
31.71Thinking Level · High
18 additional benchmarks remain in the chart above.

Standard API Pricing: GPT-6 Sol vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

GPT-6 Sol: Base price applies to <= 272000
GPT-6 Astra: Base price applies to <= 272000
ModelSupplierStandard inputStandard outputBase price applies to
GPT-6 Sol
OpenAI$2 / 1M tokens$10 / 1M tokens<= 272000
GPT-6 Astra
OpenAI$10 / 1M tokens$50 / 1M tokens<= 272000
Claude Opus 5.5
Anthropic$4 / 1M tokens$20 / 1M tokens—

Version History

How each version of the GPT-6 Sol series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGPT-6 SolCurrentGPT-5.6 SolGPT-5.5GPT-5.4
SimpleBench
Score (AVG@5)
Commonsense
73.10Standard Mode
64.80Thinking Level · Extra High
69.00Standard Mode
--
Terminal-Bench 2.1
Accuracy
Agentic Development
83.15Thinking Level · High | Tools
88.80Thinking Level · High
83.10Thinking Level · Extra High | Tools
--
Terminal-Bench 4.0
Resolution rate (%)
Agentic Development
43.94Thinking Level · High | Tools
37.27Thinking Level · High | Tools
14.60Thinking Level · Extra High | Tools
--
DeepSWE
Pass@1 (DeepSWE v1.1)
Repository Engineering
68.80Thinking Level · High | Tools
72.70Thinking Level · Extra High | Tools
67.04Thinking Level · Extra High | Tools
51.77Thinking Level · Extra High | Tools
Vals Index
跨行业任务准确率综合指数
Capability Indices
62.57Standard Mode | Tools
72.63Thinking Level · Extra High
57.41Thinking Level · Extra High | Tools
--
OSWorld 2.0
Partial score
Desktop Workflows
64.40Thinking Level · High | Tools
65.70Thinking Level · High
--
--
Agents' Last Exam
Score
Tool Orchestration
56.40Thinking Level · High | Tools
53.60Thinking Level · High
--
--
FrontierCode 1.1 Main
Score (%)
Code Generation & Editing
49.30Thinking Level · High | Tools
47.50Thinking Level · High
--
--
Program Bench
Score
Code Generation & Editing
2.00Thinking Level · High | Tools
23.00Thinking Level · High | Tools
--
--
Vibe Code Bench v1.1
Pass rate (%)
Code Generation & Editing
87.82Thinking Level · High | Tools
80.50Thinking Level · High | Tools
69.85Thinking Level · Extra High | Tools
48.47Thinking Level · Extra High | Tools
AutomationBench
Pass Rate
Office & Business
33.20Thinking Level · Extra High | Tools
45.80Thinking Level · High | Tools
--
--
CritPt
Score
Scientific Reasoning
30.86Thinking Level · High
32.30Thinking Level · High
27.10Thinking Level · Extra High
23.40Thinking Level · Extra High
17 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: SimpleBench · Commonsense

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the GPT-6 Sol Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

GPT-6 Sol: Base price applies to <= 272000
ModelSupplierStandard inputStandard outputBase price applies to
GPT-6 Sol
OpenAI$2 / 1M tokens$10 / 1M tokens<= 272000
GPT-5.6 Sol
OpenAI$4 / 1M tokens$20 / 1M tokens—
GPT-5.5
OpenAI$5 / 1M tokens$30 / 1M tokens—
GPT-5.4
OpenAI$2.5 / 1M tokens$15 / 1M tokens—
GPT-6 Sol Benchmark Results Analysis & Model Comparisons | DataLearnerAI