DataLearner 标志

DeepSeek-V4.1-FlashvsClaude Opus 5

在 13 个共同 benchmark 中,Claude Opus 5 整体领先:DeepSeek-V4.1-Flash 领先 5 项,Claude Opus 5 领先 8 项,持平 0 项,平均分差 -5.07。

DeepSeek-AI
DeepSeek-V4.1-Flash

DeepSeek-AI · 2026-09-10 · 多模态大模型

Anthropic
Claude Opus 5

Anthropic · 2026-07-24 · 推理大模型

DeepSeek-V4.1-Flash5 (38%)(62%)8 Claude Opus 5

评测分数

按能力类目分组,每组内按分差大小排列;共 13 项。

AI Agent - Tool Usage

Claude Opus 5 领先 2/3
评测项DeepSeek-V4.1-FlashClaude Opus 5分差
Terminal-Bench 4.031.207 / 16Max (With Tools)51.803 / 16Max (With Tools)-20.60
Terminal-Bench 3.0304 / 11Max (With Tools)43.301 / 11Max (With Tools)-13.30
Terminal-Bench 2.190.601 / 53Max (With Tools)89.103 / 53Max (With Tools)+1.50

Coding and Software Engineer

Claude Opus 5 领先 2/3
评测项DeepSeek-V4.1-FlashClaude Opus 5分差
Program Bench20.308 / 11Max (With Tools)375 / 11Max (With Tools)-16.70
NL2Repo-Bench65.402 / 16Max (With Tools)75.301 / 16Max (With Tools)-9.90
DeepSWE74.202 / 38Max (With Tools)744 / 38Max (With Tools)+0.20

Multimodal Understanding

Claude Opus 5 领先 3/3
评测项DeepSeek-V4.1-FlashClaude Opus 5分差
Chartography78.904 / 8Max (With Tools)842 / 8Max (With Tools)-5.10
BabyVision89.603 / 9Max (With Tools)94.101 / 9Max (With Tools)-4.50
ZeroBench Main493 / 6Max (With Tools)522 / 6Max (With Tools)-3

Agent Level Benchmark

DeepSeek-V4.1-Flash 领先 1/1
评测项DeepSeek-V4.1-FlashClaude Opus 5分差
Agents' Last Exam31.806 / 19Max (With Tools)28.607 / 19Max (With Tools)+3.20

General Knowledge

DeepSeek-V4.1-Flash 领先 1/1
评测项DeepSeek-V4.1-FlashClaude Opus 5分差
HLE63.903 / 197Max (With Tools)63.604 / 197Max (With Tools)+0.30

Productivity Knowledge

DeepSeek-V4.1-Flash 领先 1/1
评测项DeepSeek-V4.1-FlashClaude Opus 5分差
AutomationBench54.801 / 17Max (With Tools)50.303 / 17Max (With Tools)+4.50

科学与综合推理

Claude Opus 5 领先 1/1
评测项DeepSeek-V4.1-FlashClaude Opus 5分差
GPQA Diamond90.9039 / 274Max (No Tools)93.4018 / 274Max (No Tools)-2.50

规格对比

字段DeepSeek-V4.1-FlashClaude Opus 5
发布机构DeepSeek-AIAnthropic
发布时间2026-09-102026-07-24
模型类型多模态大模型推理大模型
架构MoE 架构稠密模型
参数规模5520亿暂无数据
上下文长度1M1M
最大输出384K128K

API 调用价格

价格优先使用 DataLearner 配置的 API 记录;缺失项不做推测。

价格项DeepSeek-V4.1-FlashClaude Opus 5
文本输入¥1 / 1M tokens$5 / 1M tokens
文本输出¥4 / 1M tokens$25 / 1M tokens
缓存读取¥0.02 / 1M tokens$0.5 / 1M tokens
缓存写入暂无公开价格$6.25 / 1M tokens

小结

  • DeepSeek-V4.1-Flash在以下类目领先:Agent Level Benchmark (1/1)、General Knowledge (1/1)、Productivity Knowledge (1/1)
  • Claude Opus 5在以下类目领先:AI Agent - Tool Usage (2/3)、Coding and Software Engineer (2/3)、Multimodal Understanding (3/3)、科学与综合推理 (1/1)

13 个共同 benchmark 上,Claude Opus 5 平均高出 5.07 分。

单项差距最大的 benchmark:Terminal-Bench 4.0 — DeepSeek-V4.1-Flash 31.20,Claude Opus 5 51.80(分差 -20.60)。

本页正文由结构化模型、价格与 benchmark 数据生成,不使用实时 LLM 撰写。