DataLearner logo

GPT-6 Luna Benchmark Details

GPT-6 Luna currently shows benchmark results led by AA-LCR (15 / 174, score 83), Agents' Last Exam (5 / 24, score 50.90), CritPt (43 / 204, score 19). This page also compares it with 3 competitor models and 2 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

GPT-6 Luna

Benchmark Results

Thinking
Tool usage

General Knowledge

7 evaluations
Benchmark / mode
Score
Rank/total
86.70
67 / 153
59.30
62 / 142
Vals Index
MaxTools
58.45
5 / 5
HLE
Max
39
143 / 569
19
43 / 204

Long Context

1 evaluations
Benchmark / mode
Score
Rank/total
83
15 / 174

AI Agent - Tool Usage

5 evaluations
Benchmark / mode
Score
Rank/total
73.03
89 / 199
61.48
4 / 4
53
16 / 17
OSWorld 2.0
MaxTools
52.70
11 / 13
13
48 / 95

Coding and Software Engineer

7 evaluations
Benchmark / mode
Score
Rank/total
81.65
3 / 3
DeepSWE
MaxTools
66.60
37 / 91
IOI (Vals)
MaxTools
55.56
3 / 3
55
39 / 134
42.55
2 / 2
42.40
9 / 9
0.50
15 / 15

Agent Level Benchmark

1 evaluations
Benchmark / mode
Score
Rank/total
50.90
5 / 24

Productivity Knowledge

12 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA v2
MaxTools
1367
58 / 110
AA-Briefcase
MaxTools
1299
38 / 88
MedScribe
MaxTools
83.71
2 / 3
68.52
3 / 3
58.86
2 / 3
57.65
2 / 3
49.87
4 / 5
SAGE
MaxTools
48.09
1 / 3
MedCode
MaxTools
44.69
3 / 3
30.29
3 / 4
20.70
22 / 23

Multimodal Understanding

1 evaluations
Benchmark / mode
Score
Rank/total
20
40 / 122

Other

5 evaluations
Benchmark / mode
Score
Rank/total

Other

1 evaluations
Benchmark / mode
Score
Rank/total
1
43 / 47

Math and Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total

Competitor Comparison

Benchmark scores for GPT-6 Luna compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGPT-6 LunaCurrentDeepSeek-V4.1-FlashGLM-5.3-FlashClaude Sonnet 5
AA Intelligence Index v4.3
Index (0-100)
综合评估
37.00Standard Mode | Tools
39.50Standard Mode | Tools
41.90Thinking Level · High
38.40Standard Mode | Tools
CritPt
Score
综合评估
19.00Thinking Level · High
14.30Thinking Level · High
15.40Thinking Enabled
16.90Thinking Level · High
HLE
Accuracy
综合评估
39.00Thinking Level · High
63.90Thinking Level · High | Tools
55.30Thinking Level · High | Tools
57.40Thinking Level · Extra High | Tools
AA-LCR
Accuracy
长上下文能力
83.00Thinking Level · High
84.00Thinking Level · High
--
82.00Thinking Level · High
AutomationBench-AA
Guardrail-safe objective completion rate (%)
AI Agent - 工具使用
53.00Thinking Level · High | Tools
68.90Thinking Level · High
60.40Thinking Level · High
--
Terminal-Bench 2.1
Accuracy
AI Agent - 工具使用
73.03Thinking Level · High | Tools
90.60Thinking Level · High | Tools
84.30Thinking Level · High | Tools
80.50Thinking Level · High | Tools
Terminal-Bench 4.0
Resolution rate (%)
AI Agent - 工具使用
13.00Thinking Level · High | Tools
26.80Thinking Level · High | Tools
32.80Thinking Enabled | Tools
12.42Thinking Level · High | Tools
DeepSWE
Pass@1 (DeepSWE v1.1)
编程与软件工程
66.60Thinking Level · High | Tools
74.20Thinking Level · High | Tools
63.39Thinking Level · High | Tools
54.00Deep Thinking Mode | Tools
Program Bench
Score
编程与软件工程
0.50Thinking Level · High | Tools
20.30Thinking Level · High | Tools
--
--
SciCode
Score
编程与软件工程
55.00Thinking Level · High
51.90Thinking Level · High
51.60Thinking Enabled
54.30Thinking Level · High
Agents' Last Exam
Score
Agent能力评测
50.90Thinking Level · High | Tools
31.80Thinking Level · High | Tools
26.30Thinking Level · High | Tools
--
AA-Briefcase
AA-Briefcase Elo; Rubric Pass Rate; Analytical Quality Elo; Presentation Elo
生产力知识
1299.00Thinking Level · High | Tools
1424.00Thinking Level · High | Tools
--
1355.00Thinking Level · High | Tools
4 additional benchmarks remain in the chart above.

Standard API Pricing: GPT-6 Luna vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

GPT-6 Luna: Base price applies to <= 272000
ModelSupplierStandard inputStandard outputBase price applies to
GPT-6 Luna
OpenAI$0.1 / 1M tokens$0.5 / 1M tokens<= 272000
GLM-5.3-Flash
智谱AI$0.075 / 1M tokens$0.25 / 1M tokens
Claude Sonnet 5
Anthropic$2 / 1M tokens$10 / 1M tokens

Version History

How each version of the GPT-6 Luna series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGPT-6 LunaCurrentGPT-5.6 LunaGPT-5.4 mini
AA Intelligence Index v4.3
Index (0-100)
综合评估
37.00Standard Mode | Tools
37.50Standard Mode | Tools
--
ARC-AGI-1
Score (%)
综合评估
86.70Thinking Level · High
88.00Thinking Level · High
63.67Thinking Level · Extra High | Tools
ARC-AGI-2
Score (% solved); cost per task (USD)
综合评估
59.30Thinking Level · High
59.54Thinking Level · High
18.90Thinking Level · Extra High | Tools
ARC-AGI-3 (Standard harness)
Action efficiency score(以 ARC Prize Standard harness 口径为准)
综合评估
0.19Thinking Level · Medium
0.20Thinking Level · High
--
CritPt
Score
综合评估
19.00Thinking Level · High
20.60Thinking Level · Extra High
10.00Thinking Level · Extra High
HLE
Accuracy
综合评估
39.00Thinking Level · High
39.50Thinking Level · High
41.50Thinking Level · Extra High | Tools
AA-LCR
Accuracy
长上下文能力
83.00Thinking Level · High
83.70Thinking Level · High
77.00Thinking Level · Extra High
Terminal-Bench 2.1
Accuracy
AI Agent - 工具使用
73.03Thinking Level · High | Tools
84.70Thinking Level · High
59.20Thinking Level · Extra High | Tools
Terminal-Bench 4.0
Resolution rate (%)
AI Agent - 工具使用
13.00Thinking Level · High | Tools
17.27Thinking Level · High | Tools
2.00Thinking Level · Extra High | Tools
DeepSWE
Pass@1 (DeepSWE v1.1)
编程与软件工程
66.60Thinking Level · High | Tools
67.20Thinking Level · Extra High | Tools
--
SciCode
Score
编程与软件工程
55.00Thinking Level · High
53.60Thinking Level · High
52.10Thinking Level · Extra High
Agents' Last Exam
Score
Agent能力评测
50.90Thinking Level · High | Tools
50.30Thinking Level · Extra High | Tools
--
3 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: AA Intelligence Index v4.3 · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the GPT-6 Luna Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

GPT-6 Luna: Base price applies to <= 272000
ModelSupplierStandard inputStandard outputBase price applies to
GPT-6 Luna
OpenAI$0.1 / 1M tokens$0.5 / 1M tokens<= 272000
GPT-5.6 Luna
OpenAI$0.2 / 1M tokens$1.2 / 1M tokens
GPT-5.4 mini
OpenAI$0.75 / 1M tokens$4.5 / 1M tokens