DataLearner logo

GPT-5.6 Luna Benchmark Details

GPT-5.6 Luna currently shows benchmark results led by AA-LCR (10 / 171, score 83.70), PinchBench v2 (4 / 45, score 88.67), GPQA Diamond (61 / 463, score 91.60). This page also compares it with 2 competitor models and 2 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

GPT-5.6 Luna

Benchmark Results

Thinking
Tool usage

General Knowledge

29 evaluations
Benchmark / mode
Score
Rank/total
ARC-AGI-1
LowTools
34.17
121 / 147
ARC-AGI-1
MediumTools
56.50
102 / 147
ARC-AGI-1
HighTools
76.50
77 / 147
ARC-AGI-1
Thinking Level · Max
88
53 / 147
ARC-AGI-1
Thinking Level · Extra HighTools
87.67
55 / 147
ARC-AGI-2
LowTools
5.14
104 / 136
ARC-AGI-2
MediumTools
7.36
98 / 136
ARC-AGI-2
HighTools
29.31
80 / 136
ARC-AGI-2
Thinking Level · Max
59.54
56 / 136
ARC-AGI-2
Thinking Level · Extra HighTools
47.64
67 / 136
51
21 / 28
HLE
Standard Mode
7.20
438 / 565
HLE
Low
19.80
289 / 565
HLE
Medium
25.80
242 / 565
HLE
High
33.40
185 / 565
HLE
Thinking Level · Max
39.50
137 / 565
HLE
Thinking Level · Extra High
37
158 / 565
AA Intelligence Index v4.3
Thinking Level · MaxTools
37.50
22 / 26
CritPt
Standard Mode
0.30
183 / 201
2.60
122 / 201
CritPt
Medium
4.90
99 / 201
CritPt
High
16.60
51 / 201
CritPt
Thinking Level · Max
20.60
36 / 201
CritPt
Thinking Level · Extra High
20.60
36 / 201
ARC-AGI-3 (Standard harness)
Thinking Level · Max
0.20
20 / 32
ARC-AGI-3 (Standard harness)
Thinking Level · Extra High
0.02
25 / 32

Other

6 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
63.64
365 / 463
82.32
203 / 463
85.90
150 / 463
89.20
100 / 463
GPQA Diamond
Thinking Level · Max
91.60
61 / 463
GPQA Diamond
Thinking Level · Extra High
89.50
95 / 463

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1825.80
17 / 106

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Thinking Level · Extra High
46.80
54 / 93

Text Embedding

1 evaluations
Benchmark / mode
Score
Rank/total
Context Arena
Thinking Level · Max
81.80
34 / 126

Long Context

6 evaluations
Benchmark / mode
Score
Rank/total
AA-LCR
Standard Mode
42.70
154 / 171
70
104 / 171
AA-LCR
Medium
75
85 / 171
AA-LCR
High
80.30
40 / 171
AA-LCR
Thinking Level · Max
83.70
10 / 171
AA-LCR
Thinking Level · Extra High
81.70
29 / 171

AI Agent - Tool Usage

12 evaluations
Benchmark / mode
Score
Rank/total
Terminal-Bench 2.1
Standard ModeTools
39
149 / 194
43.40
144 / 194
53.20
127 / 194
69.70
90 / 194
Terminal-Bench 2.1
Thinking Level · Max
84.70
37 / 194
Terminal-Bench 2.1
Thinking Level · MaxTools
75.70
76 / 194
Terminal-Bench 2.1
Thinking Level · Extra HighTools
77.90
70 / 194
0.50
83 / 88
2.50
65 / 88
Terminal-Bench 4.0
Thinking Level · MaxTools
17.27
34 / 88
Terminal-Bench 4.0
Thinking Level · Extra HighTools
3.50
61 / 88
Terminal-Bench-Science 0.1
Thinking Level · MaxTools
3.30
11 / 11

Coding and Software Engineer

18 evaluations
Benchmark / mode
Score
Rank/total
Text Arena (Coding)
Thinking Level · Extra High
1522.94
15 / 35
AA Coding Agent Index (historical, pre-v1.4)
Thinking Level · Extra HighTools
74.60
3 / 3
DeepSWE
LowTools
1.55
86 / 86
DeepSWE
MediumTools
11.28
85 / 86
DeepSWE
HighTools
44.25
67 / 86
DeepSWE
Thinking Level · MaxTools
67.19
29 / 86
DeepSWE
Thinking Level · Extra HighTools
67.20
28 / 86
WeirdML v2
HighTools
60.86
27 / 52
46.10
85 / 131
SciCode
Medium
46.80
82 / 131
51.60
59 / 131
SciCode
Thinking Level · Max
53.60
48 / 131
SciCode
Thinking Level · Extra High
50.50
67 / 131
16
43 / 43
CursorBench 4.0
MediumTools
22.20
42 / 43
29.40
33 / 43
CursorBench 4.0
Thinking Level · MaxTools
35.90
21 / 43
CursorBench 4.0
Thinking Level · Extra HighTools
33
27 / 43

Agent Level Benchmark

7 evaluations
Benchmark / mode
Score
Rank/total
Agents' Last Exam
Thinking Level · Extra HighTools
50.30
5 / 20
τ³-Banking
Standard ModeTools
9.70
134 / 167
τ³-Banking
LowTools
12.80
122 / 167
τ³-Banking
MediumTools
17.70
99 / 167
τ³-Banking
HighTools
25.20
77 / 167
τ³-Banking
Thinking Level · MaxTools
31.10
57 / 167
τ³-Banking
Thinking Level · Extra HighTools
28.70
67 / 167

Productivity Knowledge

13 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA v2
Standard ModeTools
1007
91 / 106
GDPval-AA v2
LowTools
1075
86 / 106
GDPval-AA v2
MediumTools
1190
73 / 106
GDPval-AA v2
HighTools
1375
52 / 106
GDPval-AA v2
Thinking Level · MaxTools
1489
31 / 106
GDPval-AA v2
Thinking Level · Extra HighTools
1428
43 / 106
AA-Briefcase
Standard ModeTools
704
79 / 84
AA-Briefcase
LowTools
744
77 / 84
AA-Briefcase
MediumTools
934
67 / 84
AA-Briefcase
HighTools
1173
49 / 84
AA-Briefcase
Thinking Level · MaxTools
1339
31 / 84
AA-Briefcase
Thinking Level · Extra HighTools
1268
39 / 84
Harvey Lab-AA
Thinking Level · MaxTools
87.90
16 / 43

Multimodal Understanding

11 evaluations
Benchmark / mode
Score
Rank/total
MMMU-Pro
Standard Mode
60.30
175 / 229
74
99 / 229
MMMU-Pro
Medium
75.80
80 / 229
77.60
66 / 229
MMMU-Pro
Thinking Level · Max
78.60
55 / 229
MMMU-Pro
Thinking Level · Extra High
78.60
55 / 229
14.20
61 / 119
GDP.pdf
Medium
15.20
58 / 119
22.20
27 / 119
GDP.pdf
Thinking Level · Max
24
19 / 119
GDP.pdf
Thinking Level · Extra High
23.80
22 / 119

Math and Reasoning

4 evaluations
Benchmark / mode
Score
Rank/total
FrontierMath v2
Standard Mode
39.65
36 / 58
41.40
35 / 58
FrontierMath v2
Thinking Level · Max
82.11
8 / 58
FrontierMath Tier 4 v2
Thinking Level · Max
60.98
9 / 42

Claw-style Agent Evaluation

1 evaluations
Benchmark / mode
Score
Rank/total
88.67
4 / 45

Competitor Comparison

Benchmark scores for GPT-5.6 Luna compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGPT-5.6 LunaCurrentGemini 3.5 FlashDeepSeek-V4-Flash
ARC-AGI-1
Score (%)
综合评估
88.00Thinking Level · High
92.50Thinking Level · High | Tools
--
ARC-AGI-2
Score (% solved); cost per task (USD)
综合评估
59.54Thinking Level · High
72.08Thinking Level · High | Tools
--
CritPt
Score
综合评估
20.60Thinking Level · Extra High
13.10Thinking Level · High
7.10Thinking Level · High
HLE
Accuracy
综合评估
39.50Thinking Level · High
42.70Thinking Level · High
51.50Thinking Level · High | Tools
GPQA Diamond
Accuracy
科学与综合推理
91.60Thinking Level · High
92.80Thinking Level · High
89.40Thinking Level · High
Creative Writing
Elo、大模型评判两两对战
写作和创作
1825.80Standard Mode
--
1555.70Standard Mode
SimpleBench
Score (AVG@5)
常识推理
46.80Thinking Level · Extra High
76.70Standard Mode
61.10Standard Mode
Context Arena
Accuracy (8 needles, 4K-128K context)
文本向量检索
81.80Thinking Level · High
77.19Thinking Level · High
69.42Thinking Enabled
AA-LCR
Accuracy
长上下文能力
83.70Thinking Level · High
74.30Thinking Level · Medium
74.30Thinking Level · High
Terminal-Bench 2.1
Accuracy
AI Agent - 工具使用
84.70Thinking Level · High
78.70Thinking Level · High | Tools
61.80Thinking Level · High | Tools
Terminal-Bench 4.0
Resolution rate (%)
AI Agent - 工具使用
17.27Thinking Level · High | Tools
6.60Thinking Level · High | Tools
3.00Thinking Level · High | Tools
DeepSWE
Pass@1 (DeepSWE v1.1)
编程与软件工程
67.20Thinking Level · Extra High | Tools
37.00Thinking Level · Medium | Tools
53.32Thinking Level · High | Tools
13 additional benchmarks remain in the chart above.

Standard API Pricing: GPT-5.6 Luna vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
GPT-5.6 Luna
OpenAI$0.2 / 1M tokens$1.2 / 1M tokens
Gemini 3.5 Flash
Google DeepMind$1.5 / 1M tokens$9 / 1M tokens
DeepSeek-V4-Flash
DeepSeek-AI$0.14 / 1M tokens$0.28 / 1M tokens

Version History

How each version of the GPT-5.6 Luna series stacks up on benchmark tests

GPT-5.6 LunaGPT-5.5GPT-5.4
Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGPT-5.6 LunaCurrentGPT-5.5GPT-5.4
AA Intelligence Index v4.3
Index (0-100)
综合评估
37.50Standard Mode | Tools
38.60Thinking Level · Extra High | Tools
--
ARC-AGI-1
Score (%)
综合评估
88.00Thinking Level · High
95.00Thinking Level · Extra High
93.67Thinking Level · Extra High
ARC-AGI-2
Score (% solved); cost per task (USD)
综合评估
59.54Thinking Level · High
85.00Thinking Level · Extra High
77.10Standard Mode
CritPt
Score
综合评估
20.60Thinking Level · Extra High
27.10Thinking Level · Extra High
23.40Thinking Level · Extra High
HLE
Accuracy
综合评估
39.50Thinking Level · High
52.20Thinking Level · High | Tools
52.10Thinking Level · Extra High | Tools
GPQA Diamond
Accuracy
科学与综合推理
91.60Thinking Level · High
94.00Thinking Level · Extra High
92.00Thinking Level · Extra High
Creative Writing
Elo、大模型评判两两对战
写作和创作
1825.80Standard Mode
1843.50Standard Mode
1835.60Standard Mode
SimpleBench
Score (AVG@5)
常识推理
46.80Thinking Level · Extra High
69.00Standard Mode
--
Context Arena
Accuracy (8 needles, 4K-128K context)
文本向量检索
81.80Thinking Level · High
94.18Thinking Level · Extra High
86.15Thinking Level · Extra High
AA-LCR
Accuracy
长上下文能力
83.70Thinking Level · High
84.30Thinking Level · Extra High
82.00Thinking Level · Extra High
Terminal-Bench 2.1
Accuracy
AI Agent - 工具使用
84.70Thinking Level · High
83.10Thinking Level · Extra High | Tools
78.30Thinking Level · Extra High | Tools
Terminal-Bench 4.0
Resolution rate (%)
AI Agent - 工具使用
17.27Thinking Level · High | Tools
14.60Thinking Level · Extra High | Tools
--
13 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: AA Intelligence Index v4.3 · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the GPT-5.6 Luna Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
GPT-5.6 Luna
OpenAI$0.2 / 1M tokens$1.2 / 1M tokens
GPT-5.5
OpenAI$5 / 1M tokens$30 / 1M tokens
GPT-5.4
OpenAI$2.5 / 1M tokens$15 / 1M tokens