GPT-6 LunavsDeepSeek-V4.1-Flash
Across 12 shared benchmarks, DeepSeek-V4.1-Flash leads overall: GPT-6 Luna wins 4, DeepSeek-V4.1-Flash wins 8, with 0 ties and an average score difference of -37.48.
GPT-6 Luna
OpenAI · 2026-09-22 · Reasoning model
DeepSeek-V4.1-Flash
DeepSeek-AI · 2026-09-10 · Multimodal model
GPT-6 Luna4 wins(33%)(67%)8 winsDeepSeek-V4.1-Flash
Benchmark scores
Grouped by capability, sorted by largest gap within each. 12 shared benchmarks.
Agentic Development
DeepSeek-V4.1-Flash 2/2| Benchmark | GPT-6 Luna | DeepSeek-V4.1-Flash | Diff |
|---|---|---|---|
| Terminal-Bench 2.1 | 73.0389 / 199Max (With Tools) | 90.604 / 199Max (With Tools) | -17.57 |
| Terminal-Bench 4.0 | 1348 / 95Max (With Tools) | 26.8028 / 95Max (With Tools) | -13.80 |
Office & Business
DeepSeek-V4.1-Flash 2/2| Benchmark | GPT-6 Luna | DeepSeek-V4.1-Flash | Diff |
|---|---|---|---|
| AA-Briefcase | 1,29938 / 88Max (With Tools) | 1,42427 / 88Max (With Tools) | -125 |
| AutomationBench | 20.7022 / 23Max (With Tools) | 54.801 / 23Max (With Tools) | -34.10 |
Code Generation & Editing
DeepSeek-V4.1-Flash 1/1| Benchmark | GPT-6 Luna | DeepSeek-V4.1-Flash | Diff |
|---|---|---|---|
| Program Bench | 0.5015 / 15Max (With Tools) | 20.309 / 15Max (With Tools) | -19.80 |
Cross-industry Work
DeepSeek-V4.1-Flash 1/1| Benchmark | GPT-6 Luna | DeepSeek-V4.1-Flash | Diff |
|---|---|---|---|
| GDPval-AA v2 | 1,36758 / 110Max (With Tools) | 1,63216 / 110Max (With Tools) | -265 |
Documents & Charts
GPT-6 Luna 1/1| Benchmark | GPT-6 Luna | DeepSeek-V4.1-Flash | Diff |
|---|---|---|---|
| GDP.pdf | 2040 / 122Max (No Tools) | 12.8070 / 122Max (No Tools) | +7.20 |
Long Reasoning
DeepSeek-V4.1-Flash 1/1| Benchmark | GPT-6 Luna | DeepSeek-V4.1-Flash | Diff |
|---|---|---|---|
| AA-LCR | 8315 / 174Max (No Tools) | 848 / 174Max (No Tools) | -1 |
Repository Engineering
DeepSeek-V4.1-Flash 1/1| Benchmark | GPT-6 Luna | DeepSeek-V4.1-Flash | Diff |
|---|---|---|---|
| DeepSWE | 66.6037 / 91Max (With Tools) | 74.202 / 91Max (With Tools) | -7.60 |
Scientific Computing
GPT-6 Luna 1/1| Benchmark | GPT-6 Luna | DeepSeek-V4.1-Flash | Diff |
|---|---|---|---|
| SciCode | 5539 / 134Max (No Tools) | 51.9061 / 134Max (No Tools) | +3.10 |
Scientific Reasoning
GPT-6 Luna 1/1| Benchmark | GPT-6 Luna | DeepSeek-V4.1-Flash | Diff |
|---|---|---|---|
| CritPt | 1943 / 204Max (No Tools) | 14.3061 / 204Max (No Tools) | +4.70 |
Tool Orchestration
GPT-6 Luna 1/1| Benchmark | GPT-6 Luna | DeepSeek-V4.1-Flash | Diff |
|---|---|---|---|
| Agents' Last Exam | 50.905 / 24Max (With Tools) | 31.808 / 24Max (With Tools) | +19.10 |
Specs
| Field | GPT-6 Luna | DeepSeek-V4.1-Flash |
|---|---|---|
| Publisher | OpenAI | DeepSeek-AI |
| Release date | 2026-09-22 | 2026-09-10 |
| Model type | Reasoning model | Multimodal model |
| Architecture | Dense | MoE |
| Parameters | Not available | 552B |
| Context length | 1.05M | 1M |
| Max output | 128K | 384K |
API pricing
Prices use DataLearner records when available; missing fields are not inferred.
| Item | GPT-6 Luna | DeepSeek-V4.1-Flash |
|---|---|---|
| Text input | $0.1 / 1M tokens | ¥1 / 1M tokens |
| Text output | $0.5 / 1M tokens | ¥4 / 1M tokens |
| Cache read | $0.01 / 1M tokens | ¥0.02 / 1M tokens |
| Cache write | $0.125 / 1M tokens | Not public |
Summary
- GPT-6 Lunaleads in:Documents & Charts (1/1), Scientific Computing (1/1), Scientific Reasoning (1/1), Tool Orchestration (1/1)
- DeepSeek-V4.1-Flashleads in:Agentic Development (2/2), Office & Business (2/2), Code Generation & Editing (1/1), Cross-industry Work (1/1), Long Reasoning (1/1), Repository Engineering (1/1)
On average across the 12 shared benchmarks, DeepSeek-V4.1-Flash scores 37.48 higher.
Largest single-benchmark gap: GDPval-AA v2 — GPT-6 Luna 1,367 vs DeepSeek-V4.1-Flash 1,632 (-265).
Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.