DataLearner logo

GPT-6 AstravsGPT-5.4

Across 8 shared benchmarks, GPT-6 Astra leads overall: GPT-6 Astra wins 8, GPT-5.4 wins 0, with 0 ties and an average score difference of +23.91.

OpenAI
GPT-6 Astra

OpenAI · 2026-09-03 · Reasoning model

OpenAI
GPT-5.4

OpenAI · 2026-03-05 · Multimodal model

GPT-6 Astra8 wins(100%)(0%)0 winsGPT-5.4

Benchmark scores

Grouped by capability, sorted by largest gap within each. 8 shared benchmarks.

General Knowledge

GPT-6 Astra 4/4
BenchmarkGPT-6 AstraGPT-5.4Diff
ARC-AGI-362.701 / 16Max (No Tools)014 / 16Thinking High (No Tools)+62.70
ARC-AGI-2951 / 85Max (No Tools)77.1023 / 85Normal (No Tools)+17.90
HLE57.2012 / 190Max (With Tools)52.1028 / 190Extra-High (With Tools)+5.10
ARC-AGI-198.501 / 91Max (No Tools)93.7021 / 91Normal (No Tools)+4.80

AI Agent - Information Search

GPT-6 Astra 1/1
BenchmarkGPT-6 AstraGPT-5.4Diff
BrowseComp91.501 / 56Max (With Tools)82.7016 / 56Extra-High (With Tools)+8.80

Coding and Software Engineer

GPT-6 Astra 1/1
BenchmarkGPT-6 AstraGPT-5.4Diff
DeepSWE74.102 / 35Max (With Tools)5227 / 35Extra-High (With Tools)+22.10

General Evaluation

GPT-6 Astra 1/1
BenchmarkGPT-6 AstraGPT-5.4Diff
GPQA Diamond961 / 271Max (No Tools)74.75174 / 271Normal (No Tools)+21.25

Math and Reasoning

GPT-6 Astra 1/1
BenchmarkGPT-6 AstraGPT-5.4Diff
FrontierMath Tier 4 v297.601 / 41Max (No Tools)4911 / 41Extra-High (No Tools)+48.60

Specs

FieldGPT-6 AstraGPT-5.4
PublisherOpenAIOpenAI
Release date2026-09-032026-03-05
Model typeReasoning modelMultimodal model
ArchitectureDenseDense
ParametersNot availableNot available
Context length1.05M1M
Max output128K125K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGPT-6 AstraGPT-5.4
Text input$10 / 1M tokens$2.5 / 1M tokens
Text output$50 / 1M tokens$15 / 1M tokens
Cache read$1 / 1M tokens$0.25 / 1M tokens
Cache write$12.5 / 1M tokensNot public

Summary

  • GPT-6 Astraleads in:General Knowledge (4/4), AI Agent - Information Search (1/1), Coding and Software Engineer (1/1), General Evaluation (1/1), Math and Reasoning (1/1)

On average across the 8 shared benchmarks, GPT-6 Astra scores 23.91 higher.

Largest single-benchmark gap: ARC-AGI-3 — GPT-6 Astra 62.70 vs GPT-5.4 0 (+62.70).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.