GPT-5.6 Luna Benchmark Details
GPT-5.6 Luna currently shows benchmark results led by AA-LCR (10 / 171, score 83.70), PinchBench v2 (4 / 45, score 88.67), GPQA Diamond (61 / 463, score 91.60). This page also compares it with 2 competitor models and 2 predecessor or same-series models, including performance and pricing views when available.
Benchmark Results
Benchmark Results
General Knowledge
29 evaluationsOther
6 evaluationsWriting and Creative Capabilities
1 evaluationsLong Context
6 evaluationsAI Agent - Tool Usage
12 evaluationsCoding and Software Engineer
18 evaluationsAgent Level Benchmark
7 evaluationsProductivity Knowledge
13 evaluationsMultimodal Understanding
11 evaluationsMath and Reasoning
4 evaluationsCompetitor Comparison
Benchmark scores for GPT-5.6 Luna compared against top models in its class
12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.
| Benchmark | GPT-5.6 LunaCurrent | Gemini 3.5 Flash | DeepSeek-V4-Flash |
|---|---|---|---|
88.00Thinking Level · High | 92.50Thinking Level · High | Tools | -- | |
59.54Thinking Level · High | 72.08Thinking Level · High | Tools | -- | |
20.60Thinking Level · Extra High | 13.10Thinking Level · High | 7.10Thinking Level · High | |
39.50Thinking Level · High | 42.70Thinking Level · High | 51.50Thinking Level · High | Tools | |
91.60Thinking Level · High | 92.80Thinking Level · High | 89.40Thinking Level · High | |
1825.80Standard Mode | -- | 1555.70Standard Mode | |
46.80Thinking Level · Extra High | 76.70Standard Mode | 61.10Standard Mode | |
81.80Thinking Level · High | 77.19Thinking Level · High | 69.42Thinking Enabled | |
83.70Thinking Level · High | 74.30Thinking Level · Medium | 74.30Thinking Level · High | |
84.70Thinking Level · High | 78.70Thinking Level · High | Tools | 61.80Thinking Level · High | Tools | |
17.27Thinking Level · High | Tools | 6.60Thinking Level · High | Tools | 3.00Thinking Level · High | Tools | |
67.20Thinking Level · Extra High | Tools | 37.00Thinking Level · Medium | Tools | 53.32Thinking Level · High | Tools |
Standard API Pricing: GPT-5.6 Luna vs. Peer Models
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
GPT-5.6 Luna | OpenAI | $0.2 / 1M tokens | $1.2 / 1M tokens | — |
Gemini 3.5 Flash | Google DeepMind | $1.5 / 1M tokens | $9 / 1M tokens | — |
DeepSeek-V4-Flash | DeepSeek-AI | $0.14 / 1M tokens | $0.28 / 1M tokens | — |
Version History
How each version of the GPT-5.6 Luna series stacks up on benchmark tests
12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.
| Benchmark | GPT-5.6 LunaCurrent | GPT-5.5 | GPT-5.4 |
|---|---|---|---|
37.50Standard Mode | Tools | 38.60Thinking Level · Extra High | Tools | -- | |
88.00Thinking Level · High | 95.00Thinking Level · Extra High | 93.67Thinking Level · Extra High | |
59.54Thinking Level · High | 85.00Thinking Level · Extra High | 77.10Standard Mode | |
20.60Thinking Level · Extra High | 27.10Thinking Level · Extra High | 23.40Thinking Level · Extra High | |
39.50Thinking Level · High | 52.20Thinking Level · High | Tools | 52.10Thinking Level · Extra High | Tools | |
91.60Thinking Level · High | 94.00Thinking Level · Extra High | 92.00Thinking Level · Extra High | |
1825.80Standard Mode | 1843.50Standard Mode | 1835.60Standard Mode | |
46.80Thinking Level · Extra High | 69.00Standard Mode | -- | |
81.80Thinking Level · High | 94.18Thinking Level · Extra High | 86.15Thinking Level · Extra High | |
83.70Thinking Level · High | 84.30Thinking Level · Extra High | 82.00Thinking Level · Extra High | |
84.70Thinking Level · High | 83.10Thinking Level · Extra High | Tools | 78.30Thinking Level · Extra High | Tools | |
17.27Thinking Level · High | Tools | 14.60Thinking Level · Extra High | Tools | -- |
Single-Benchmark Version Trend
Viewing: AA Intelligence Index v4.3 · 综合评估
Standard API Pricing Across the GPT-5.6 Luna Series
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
GPT-5.6 Luna | OpenAI | $0.2 / 1M tokens | $1.2 / 1M tokens | — |
GPT-5.5 | OpenAI | $5 / 1M tokens | $30 / 1M tokens | — |
GPT-5.4 | OpenAI | $2.5 / 1M tokens | $15 / 1M tokens | — |