DataLearner 标志

DeepSeek-V4.1-FlashvsDeepSeek V4 Flash 0731

在 12 个同模式 benchmark 中,DeepSeek-V4.1-Flash 整体领先:DeepSeek-V4.1-Flash 领先 10 项,DeepSeek V4 Flash 0731 领先 2 项,持平 0 项,平均分差 +9.20。另有 1 项测试模式不同,仅供参考。

DeepSeek-AI
DeepSeek-V4.1-Flash

DeepSeek-AI · 2026-09-10 · 多模态大模型

DeepSeek-AI
DeepSeek V4 Flash 0731

DeepSeek-AI · 2026-07-31 · 推理大模型

DeepSeek-V4.1-Flash10 项(83%)(17%)2 项DeepSeek V4 Flash 0731

评测分数

按能力类目分组,每组内按分差大小排列;共 12 项。

仓库修复与多文件工程

DeepSeek-V4.1-Flash 领先 2/2
评测项DeepSeek-V4.1-FlashDeepSeek V4 Flash 0731分差
DeepSWE74.204 / 97Max (With Tools)54.4059 / 97Max (With Tools)+19.80
NL2Repo-Bench65.402 / 18Max (With Tools)54.2012 / 18Max (With Tools)+11.20

办公与文档流程

DeepSeek-V4.1-Flash 领先 1/1
评测项DeepSeek-V4.1-FlashDeepSeek V4 Flash 0731分差
AutomationBench54.801 / 28Max (With Tools)25.1025 / 28Max (With Tools)+29.70

抽象归纳与泛化

DeepSeek V4 Flash 0731 领先 1/1
评测项DeepSeek-V4.1-FlashDeepSeek V4 Flash 0731分差
ARC-AGI-17.83214 / 221Normal (No Tools)11.83209 / 221Normal (No Tools)-4

法律

DeepSeek V4 Flash 0731 领先 1/1
评测项DeepSeek-V4.1-FlashDeepSeek V4 Flash 0731分差
Harvey's Legal Agent Benchmark6.6720 / 50Thinking High (With Tools)8.3316 / 50Thinking High (With Tools)-1.66

漏洞发现与分析

DeepSeek-V4.1-Flash 领先 1/1
评测项DeepSeek-V4.1-FlashDeepSeek V4 Flash 0731分差
CyberGym88.103 / 13Max (With Tools)76.7010 / 13Max (With Tools)+11.40

生物与基因

DeepSeek-V4.1-Flash 领先 1/1
评测项DeepSeek-V4.1-FlashDeepSeek V4 Flash 0731分差
BioMysteryBench67.7813 / 21Thinking High (With Tools)64.4416 / 21Thinking High (With Tools)+3.34

程序生成与编辑

DeepSeek-V4.1-Flash 领先 1/1
评测项DeepSeek-V4.1-FlashDeepSeek V4 Flash 0731分差
Vibe Code Bench v1.184.7413 / 65Thinking High (With Tools)74.7426 / 65Thinking High (With Tools)+10

维护、优化与工程管理

DeepSeek-V4.1-Flash 领先 1/1
评测项DeepSeek-V4.1-FlashDeepSeek V4 Flash 0731分差
Code Migration45.6214 / 48Thinking High (With Tools)38.6325 / 48Thinking High (With Tools)+6.99

综合知识考试

DeepSeek-V4.1-Flash 领先 1/1
评测项DeepSeek-V4.1-FlashDeepSeek V4 Flash 0731分差
HLE63.905 / 240Max (With Tools)51.5037 / 240Max (With Tools)+12.40

自主开发与终端任务

DeepSeek-V4.1-Flash 领先 1/1
评测项DeepSeek-V4.1-FlashDeepSeek V4 Flash 0731分差
Terminal-Bench 2.190.602 / 61Max (With Tools)82.7026 / 61Max (With Tools)+7.90

跨能力聚合指数

DeepSeek-V4.1-Flash 领先 1/1
评测项DeepSeek-V4.1-FlashDeepSeek V4 Flash 0731分差
Vals Index v2.151.3222 / 37Thinking High (With Tools)47.9630 / 37Thinking High (With Tools)+3.36

测试模式不同的成绩

共 1 项。两款模型的公开成绩来自不同测试模式,分数不直接可比,不计入胜负和平均分差。

评测项DeepSeek-V4.1-FlashDeepSeek V4 Flash 0731
LiveBench81.11Max (No Tools)74.17Reported best (effort unspecified)

规格对比

字段DeepSeek-V4.1-FlashDeepSeek V4 Flash 0731
发布机构DeepSeek-AIDeepSeek-AI
发布时间2026-09-102026-07-31
模型类型多模态大模型推理大模型
架构MoE 架构稠密模型
参数规模5520亿暂无数据
上下文长度1M暂无数据
最大输出384K暂无数据

API 调用价格

价格优先使用 DataLearner 配置的 API 记录;缺失项不做推测。

价格项DeepSeek-V4.1-FlashDeepSeek V4 Flash 0731
文本输入¥1 / 1M tokens暂无公开价格
文本输出¥4 / 1M tokens暂无公开价格
缓存读取¥0.02 / 1M tokens暂无公开价格

部分模型公开价格不完整,缺失字段按"暂无公开价格"展示。

小结

  • DeepSeek-V4.1-Flash在以下类目领先:仓库修复与多文件工程 (2/2)、办公与文档流程 (1/1)、漏洞发现与分析 (1/1)、生物与基因 (1/1)、程序生成与编辑 (1/1)、维护、优化与工程管理 (1/1)、综合知识考试 (1/1)、自主开发与终端任务 (1/1)、跨能力聚合指数 (1/1)
  • DeepSeek V4 Flash 0731在以下类目领先:抽象归纳与泛化 (1/1)、法律 (1/1)

12 个同模式 benchmark 上,DeepSeek-V4.1-Flash 平均高出 9.20 分。

单项差距最大的 benchmark:AutomationBench — DeepSeek-V4.1-Flash 54.80,DeepSeek V4 Flash 0731 25.10(分差 +29.70)。

本页正文由结构化模型、价格与 benchmark 数据生成,不使用实时 LLM 撰写。