DataLearner logo

GPT-6 Astra Benchmark Details

GPT-6 Astra currently shows benchmark results led by GPQA Diamond (1 / 271, score 96), ARC-AGI-1 (1 / 91, score 98.50), ARC-AGI-2 (1 / 85, score 95). This page also tracks comparisons against 3 predecessor or same-series models.

Benchmark Results

GPT-6 Astra

Benchmark Results

Thinking
Tool usage

General Knowledge

8 evaluations
Benchmark / mode
Score
Rank/total
98.50
1 / 91
95
1 / 85
62.70
1 / 16
61.20
7 / 28
HLE
MaxTools
57.20
12 / 190
2.40
1 / 1
0
1 / 1

Other

5 evaluations
Benchmark / mode
Score
Rank/total

AI Agent - Information Search

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
MaxTools
91.50
1 / 56

Coding and Software Engineer

7 evaluations
Benchmark / mode
Score
Rank/total
85.40
1 / 1
DeepSWE
MaxTools
74.10
2 / 35
67
4 / 4
64.50
1 / 6
63.90
1 / 1
53.30
1 / 1

AI Agent - Tool Usage

10 evaluations
Benchmark / mode
Score
Rank/total
ExploitBench
MaxTools
100
1 / 2
100
1 / 1
OSWorld 2.0
MaxTools
72.60
2 / 9
64.60
1 / 11
57.90
1 / 13
42.40
1 / 1
41.40
5 / 14
0
1 / 1

Agent Level Benchmark

1 evaluations
Benchmark / mode
Score
Rank/total
59.30
1 / 15

Math and Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total

Productivity Knowledge

3 evaluations
Benchmark / mode
Score
Rank/total
BenchCAD
MaxTools
95.90
1 / 1
50
1 / 1
40.90
1 / 1

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total

Other

1 evaluations
Benchmark / mode
Score
Rank/total
4.20
1 / 1

Other

2 evaluations
Benchmark / mode
Score
Rank/total

Version History

How each version of the GPT-6 Astra series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGPT-6 AstraCurrentGPT-5.6 SolGPT-5.5GPT-5.4
61.20Thinking Level · High
61.00Thinking Level · High | Tools
--
--
ARC-AGI-1
综合评估
98.50Thinking Level · High
97.50Thinking Level · Extra High
95.00Thinking Level · Extra High
93.70Thinking Level · Extra High
ARC-AGI-2
综合评估
95.00Thinking Level · High
92.50Thinking Level · High
85.00Thinking Level · Extra High
77.10Standard Mode
ARC-AGI-3
综合评估
62.70Thinking Level · High
7.80Thinking Level · High
0.00Thinking Level · High
0.00Thinking Level · High
HLE
综合评估
57.20Thinking Level · High | Tools
49.50Thinking Level · High
52.20Thinking Level · High | Tools
52.10Thinking Level · Extra High | Tools
GPQA Diamond
科学与综合推理
96.00Thinking Level · High
93.50Thinking Level · High
94.00Thinking Level · Extra High
92.80Thinking Level · Extra High
BrowseComp
AI Agent - 信息收集
91.50Thinking Level · High | Tools
--
84.40Thinking Level · High | Tools
82.70Thinking Level · Extra High | Tools
AA Coding Agent Index
编程与软件工程
67.00Thinking Level · High | Tools
80.00Thinking Level · Extra High | Tools
--
--
DeepSWE
编程与软件工程
74.10Thinking Level · High | Tools
72.70Thinking Level · Extra High | Tools
67.00Thinking Level · Extra High | Tools
52.00Thinking Level · Extra High | Tools
FrontierCode 1.1
编程与软件工程
64.50Thinking Level · High | Tools
60.60Thinking Level · High | Tools
--
--
OSWorld 2.0
AI Agent - 工具使用
72.60Thinking Level · High | Tools
62.60Thinking Level · Extra High | Tools
--
--
Terminal-Bench 4.0
AI Agent - 工具使用
57.90Thinking Level · High | Tools
37.27Thinking Level · High | Tools
--
--
3 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: AA Intelligence Index · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the GPT-6 Astra Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

GPT-6 Astra: Base price applies to <= 272000
ModelSupplierStandard inputStandard outputBase price applies to
GPT-6 Astra
OpenAI$10 / 1M tokens$50 / 1M tokens<= 272000
GPT-5.6 Sol
OpenAI$4 / 1M tokens$20 / 1M tokens
GPT-5.5
OpenAI$5 / 1M tokens$30 / 1M tokens
GPT-5.4
OpenAI$2.5 / 1M tokens$15 / 1M tokens