GPT-6 Astra Benchmark Details
GPT-6 Astra currently shows benchmark results led by GPQA Diamond (1 / 271, score 96), ARC-AGI-1 (1 / 91, score 98.50), ARC-AGI-2 (1 / 85, score 95). This page also tracks comparisons against 3 predecessor or same-series models.
Benchmark Results
Benchmark Results
General Knowledge
8 evaluationsOther
5 evaluationsCoding and Software Engineer
7 evaluationsAI Agent - Tool Usage
10 evaluationsProductivity Knowledge
3 evaluationsWriting and Creative Capabilities
1 evaluationsOther
2 evaluationsVersion History
How each version of the GPT-6 Astra series stacks up on benchmark tests
12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.
| Benchmark | GPT-6 AstraCurrent | GPT-5.6 Sol | GPT-5.5 | GPT-5.4 |
|---|---|---|---|---|
61.20Thinking Level · High | 61.00Thinking Level · High | Tools | -- | -- | |
ARC-AGI-1 综合评估 | 98.50Thinking Level · High | 97.50Thinking Level · Extra High | 95.00Thinking Level · Extra High | 93.70Thinking Level · Extra High |
ARC-AGI-2 综合评估 | 95.00Thinking Level · High | 92.50Thinking Level · High | 85.00Thinking Level · Extra High | 77.10Standard Mode |
ARC-AGI-3 综合评估 | 62.70Thinking Level · High | 7.80Thinking Level · High | 0.00Thinking Level · High | 0.00Thinking Level · High |
HLE 综合评估 | 57.20Thinking Level · High | Tools | 49.50Thinking Level · High | 52.20Thinking Level · High | Tools | 52.10Thinking Level · Extra High | Tools |
GPQA Diamond 科学与综合推理 | 96.00Thinking Level · High | 93.50Thinking Level · High | 94.00Thinking Level · Extra High | 92.80Thinking Level · Extra High |
BrowseComp AI Agent - 信息收集 | 91.50Thinking Level · High | Tools | -- | 84.40Thinking Level · High | Tools | 82.70Thinking Level · Extra High | Tools |
AA Coding Agent Index 编程与软件工程 | 67.00Thinking Level · High | Tools | 80.00Thinking Level · Extra High | Tools | -- | -- |
DeepSWE 编程与软件工程 | 74.10Thinking Level · High | Tools | 72.70Thinking Level · Extra High | Tools | 67.00Thinking Level · Extra High | Tools | 52.00Thinking Level · Extra High | Tools |
FrontierCode 1.1 编程与软件工程 | 64.50Thinking Level · High | Tools | 60.60Thinking Level · High | Tools | -- | -- |
OSWorld 2.0 AI Agent - 工具使用 | 72.60Thinking Level · High | Tools | 62.60Thinking Level · Extra High | Tools | -- | -- |
Terminal-Bench 4.0 AI Agent - 工具使用 | 57.90Thinking Level · High | Tools | 37.27Thinking Level · High | Tools | -- | -- |
Single-Benchmark Version Trend
Viewing: AA Intelligence Index · 综合评估
Standard API Pricing Across the GPT-6 Astra Series
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens
When a context threshold exists, the charted base price only applies within these limits:
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
GPT-6 Astra | OpenAI | $10 / 1M tokens | $50 / 1M tokens | <= 272000 |
GPT-5.6 Sol | OpenAI | $4 / 1M tokens | $20 / 1M tokens | — |
GPT-5.5 | OpenAI | $5 / 1M tokens | $30 / 1M tokens | — |
GPT-5.4 | OpenAI | $2.5 / 1M tokens | $15 / 1M tokens | — |