GPT-6 Astra leads on the shared-benchmark average
Ahead on 5 of 8 shared benchmarks, averaging 1.5 points higher
Summarised only from the 8 percentage-scale benchmarks scored by every selected model; details are below.
“Best available” takes each model’s highest recorded non-parallel mode per benchmark, so it may combine modes into a virtual configuration that does not exist. Read it with the mode breakdown.

Claude Fable 5.1
Anthropic
Benchmark-by-benchmark comparison. Changing the thinking mode or tool filters updates the chart and table below.
“Best available” picks the highest non-parallel mode separately for each benchmark. The resulting series can combine several reasoning levels and is not one reproducible runtime configuration. Choose a mode filter to compare like-for-like runs.
Data sourced primarily from official releases (GitHub, Hugging Face, papers), then benchmark leaderboards, then third-party evaluators. Learn about our data methodology
Every model and runtime mode, benchmark by benchmark. Values are comparable along a row, not between different benchmarks.
8 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.
| Benchmark | Claude Fable 5.1 | GPT-6 Astra |
|---|---|---|
66.00Thinking Level · High | Tools | 61.20Thinking Level · High | |
ARC-AGI-1 综合评估 | 97.50Thinking Level · High | 98.50Thinking Level · High |
ARC-AGI-2 综合评估 | 90.00Thinking Level · Extra High | 95.00Thinking Level · High |
HLE 综合评估 | 65.00Thinking Level · High | Tools | 57.20Thinking Level · High | Tools |
AutomationBench AI Agent - 工具使用 | 31.40Thinking Level · High | Tools | 41.40Thinking Level · High | Tools |
OSWorld 2.0 AI Agent - 工具使用 | 77.90Thinking Level · High | Tools | 72.60Thinking Level · High | Tools |
Terminal-Bench 4.0 AI Agent - 工具使用 | 55.80Thinking Level · High | Tools | 57.90Thinking Level · High | Tools |
Terminal-Bench-Science 0.1 AI Agent - 工具使用 | 52.60Thinking Level · High | Tools | 64.60Thinking Level · High | Tools |
Official list prices per model API, split by input and output. Unit: USD per 1M tokens.
Architecture, licensing and API modalities. "Not provided" means the field is missing from our database.
| Features & specs | Claude Fable 5.1Anthropic | GPT-6 AstraOpenAI |
|---|---|---|
Core specsRelease | 2026-09-01 | 2026-09-03 |
Context length | 1M | 1.05M |
Max output length | 128,000 tokens | 128,000 tokens |
Architecture | Undisclosed | Undisclosed |
Runtime modes | 自适应思考 | 低中高极高最高 |
Availability & licensingCode availability | Not public | Not public |
Weight availability | Not public | Not public |
Use & commercial terms | Official service only; subject to provider terms | Official service only; subject to provider terms |
API modality supportText Input/Output | Input:YesOutput:Yes | Input:YesOutput:Yes |
Image Input/Output | Input:YesOutput:No | Input:YesOutput:No |
ResourcesPaper / report | Claude Fable 5.1 and Claude Mythos 5.1 System Card | GPT-6 Astra Model | OpenAI API |

GPT-6 Astra
OpenAI