DataLearner logo

GPT-6 LunavsDeepSeek-V4.1-Flash

Across 12 shared benchmarks, DeepSeek-V4.1-Flash leads overall: GPT-6 Luna wins 4, DeepSeek-V4.1-Flash wins 8, with 0 ties and an average score difference of -37.48.

OpenAI
GPT-6 Luna

OpenAI · 2026-09-22 · Reasoning model

DeepSeek-AI
DeepSeek-V4.1-Flash

DeepSeek-AI · 2026-09-10 · Multimodal model

GPT-6 Luna4 wins(33%)(67%)8 winsDeepSeek-V4.1-Flash

Benchmark scores

Grouped by capability, sorted by largest gap within each. 12 shared benchmarks.

Agentic Development

DeepSeek-V4.1-Flash 2/2
BenchmarkGPT-6 LunaDeepSeek-V4.1-FlashDiff
Terminal-Bench 2.173.0389 / 199Max (With Tools)90.604 / 199Max (With Tools)-17.57
Terminal-Bench 4.01348 / 95Max (With Tools)26.8028 / 95Max (With Tools)-13.80

Office & Business

DeepSeek-V4.1-Flash 2/2
BenchmarkGPT-6 LunaDeepSeek-V4.1-FlashDiff
AA-Briefcase1,29938 / 88Max (With Tools)1,42427 / 88Max (With Tools)-125
AutomationBench20.7022 / 23Max (With Tools)54.801 / 23Max (With Tools)-34.10

Code Generation & Editing

DeepSeek-V4.1-Flash 1/1
BenchmarkGPT-6 LunaDeepSeek-V4.1-FlashDiff
Program Bench0.5015 / 15Max (With Tools)20.309 / 15Max (With Tools)-19.80

Cross-industry Work

DeepSeek-V4.1-Flash 1/1
BenchmarkGPT-6 LunaDeepSeek-V4.1-FlashDiff
GDPval-AA v21,36758 / 110Max (With Tools)1,63216 / 110Max (With Tools)-265

Documents & Charts

GPT-6 Luna 1/1
BenchmarkGPT-6 LunaDeepSeek-V4.1-FlashDiff
GDP.pdf2040 / 122Max (No Tools)12.8070 / 122Max (No Tools)+7.20

Long Reasoning

DeepSeek-V4.1-Flash 1/1
BenchmarkGPT-6 LunaDeepSeek-V4.1-FlashDiff
AA-LCR8315 / 174Max (No Tools)848 / 174Max (No Tools)-1

Repository Engineering

DeepSeek-V4.1-Flash 1/1
BenchmarkGPT-6 LunaDeepSeek-V4.1-FlashDiff
DeepSWE66.6037 / 91Max (With Tools)74.202 / 91Max (With Tools)-7.60

Scientific Computing

GPT-6 Luna 1/1
BenchmarkGPT-6 LunaDeepSeek-V4.1-FlashDiff
SciCode5539 / 134Max (No Tools)51.9061 / 134Max (No Tools)+3.10

Scientific Reasoning

GPT-6 Luna 1/1
BenchmarkGPT-6 LunaDeepSeek-V4.1-FlashDiff
CritPt1943 / 204Max (No Tools)14.3061 / 204Max (No Tools)+4.70

Tool Orchestration

GPT-6 Luna 1/1
BenchmarkGPT-6 LunaDeepSeek-V4.1-FlashDiff
Agents' Last Exam50.905 / 24Max (With Tools)31.808 / 24Max (With Tools)+19.10

Specs

FieldGPT-6 LunaDeepSeek-V4.1-Flash
PublisherOpenAIDeepSeek-AI
Release date2026-09-222026-09-10
Model typeReasoning modelMultimodal model
ArchitectureDenseMoE
ParametersNot available552B
Context length1.05M1M
Max output128K384K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGPT-6 LunaDeepSeek-V4.1-Flash
Text input$0.1 / 1M tokens¥1 / 1M tokens
Text output$0.5 / 1M tokens¥4 / 1M tokens
Cache read$0.01 / 1M tokens¥0.02 / 1M tokens
Cache write$0.125 / 1M tokensNot public

Summary

  • GPT-6 Lunaleads in:Documents & Charts (1/1), Scientific Computing (1/1), Scientific Reasoning (1/1), Tool Orchestration (1/1)
  • DeepSeek-V4.1-Flashleads in:Agentic Development (2/2), Office & Business (2/2), Code Generation & Editing (1/1), Cross-industry Work (1/1), Long Reasoning (1/1), Repository Engineering (1/1)

On average across the 12 shared benchmarks, DeepSeek-V4.1-Flash scores 37.48 higher.

Largest single-benchmark gap: GDPval-AA v2 — GPT-6 Luna 1,367 vs DeepSeek-V4.1-Flash 1,632 (-265).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.