DataLearner logo

GPT-6 LunavsGPT-5.6 Luna

Across 4 shared benchmarks, GPT-5.6 Luna leads overall: GPT-6 Luna wins 1, GPT-5.6 Luna wins 3, with 0 ties and an average score difference of -2.04.

OpenAI
GPT-6 Luna

OpenAI · 2026-09-22 · Reasoning model

OpenAI
GPT-5.6 Luna

OpenAI · 2026-06-26 · Reasoning model

GPT-6 Luna1 win(25%)(75%)3 winsGPT-5.6 Luna

Benchmark scores

Grouped by capability, sorted by largest gap within each. 4 shared benchmarks.

AI Agent - Tool Usage

GPT-5.6 Luna 1/1
BenchmarkGPT-6 LunaGPT-5.6 LunaDiff
Terminal-Bench 4.01348 / 95Max (With Tools)17.2740 / 95Max (With Tools)-4.27

Coding and Software Engineer

GPT-6 Luna 1/1
BenchmarkGPT-6 LunaGPT-5.6 LunaDiff
SciCode5539 / 134Max (No Tools)53.6051 / 134Max (No Tools)+1.40

General Knowledge

GPT-5.6 Luna 1/1
BenchmarkGPT-6 LunaGPT-5.6 LunaDiff
ARC-AGI-186.7067 / 153Max (No Tools)8858 / 153Max (No Tools)-1.30

Multimodal Understanding

GPT-5.6 Luna 1/1
BenchmarkGPT-6 LunaGPT-5.6 LunaDiff
GDP.pdf2040 / 122Max (No Tools)2421 / 122Max (No Tools)-4

Specs

FieldGPT-6 LunaGPT-5.6 Luna
PublisherOpenAIOpenAI
Release date2026-09-222026-06-26
Model typeReasoning modelReasoning model
ArchitectureDenseDense
ParametersNot availableNot available
Context length1.05M1.05M
Max output128K128K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGPT-6 LunaGPT-5.6 Luna
Text input$0.1 / 1M tokens$0.2 / 1M tokens
Text output$0.5 / 1M tokens$1.2 / 1M tokens
Cache read$0.01 / 1M tokens$0.02 / 1M tokens
Cache write$0.125 / 1M tokens$0.25 / 1M tokens

Summary

  • GPT-6 Lunaleads in:Coding and Software Engineer (1/1)
  • GPT-5.6 Lunaleads in:AI Agent - Tool Usage (1/1), General Knowledge (1/1), Multimodal Understanding (1/1)

On average across the 4 shared benchmarks, GPT-5.6 Luna scores 2.04 higher.

Largest single-benchmark gap: Terminal-Bench 4.0 — GPT-6 Luna 13 vs GPT-5.6 Luna 17.27 (-4.27).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.