DataLearner logo

DeepSeek-V4.1-FlashvsKimi K3

Across 15 shared benchmarks, DeepSeek-V4.1-Flash leads overall: DeepSeek-V4.1-Flash wins 13, Kimi K3 wins 1, with 1 ties and an average score difference of +6.35.

DeepSeek-AI
DeepSeek-V4.1-Flash

DeepSeek-AI · 2026-09-10 · Multimodal model

Moonshot AI
Kimi K3

Moonshot AI · 2026-07-16 · Reasoning model

DeepSeek-V4.1-Flash13 wins(87%)Ties1(7%)1 winKimi K3

Benchmark scores

Grouped by capability, sorted by largest gap within each. 15 shared benchmarks.

AI Agent - Tool Usage

DeepSeek-V4.1-Flash 4/4
BenchmarkDeepSeek-V4.1-FlashKimi K3Diff
Terminal-Bench 4.031.207 / 16Max (With Tools)12.6013 / 16Max (With Tools)+18.60
Terminal-Bench 3.0304 / 11Max (With Tools)17.708 / 11Max (With Tools)+12.30
CyberGym88.101 / 8Max (With Tools)805 / 8Max (With Tools)+8.10
Terminal-Bench 2.190.601 / 53Max (With Tools)88.307 / 53Max (With Tools)+2.30

Coding and Software Engineer

DeepSeek-V4.1-Flash 3/3
BenchmarkDeepSeek-V4.1-FlashKimi K3Diff
NL2Repo-Bench65.402 / 16Max (With Tools)586 / 16Max (With Tools)+7.40
DeepSWE74.202 / 38Max (With Tools)67.5011 / 38Max (With Tools)+6.70
Program Bench20.308 / 11Max (With Tools)17.5010 / 11Max (With Tools)+2.80

Multimodal Understanding

DeepSeek-V4.1-Flash 3/3
BenchmarkDeepSeek-V4.1-FlashKimi K3Diff
Chartography78.904 / 8Max (With Tools)68.106 / 8Max (With Tools)+10.80
ZeroBench Main493 / 6Max (With Tools)414 / 6Max (With Tools)+8
BabyVision89.603 / 9Max (With Tools)85.705 / 9Max (With Tools)+3.90

Agent Level Benchmark

DeepSeek-V4.1-Flash 1/1
BenchmarkDeepSeek-V4.1-FlashKimi K3Diff
Agents' Last Exam31.806 / 19Max (With Tools)27.609 / 19Max (With Tools)+4.20

General Evaluation

Kimi K3 1/1
BenchmarkDeepSeek-V4.1-FlashKimi K3Diff
GPQA Diamond90.9039 / 274Max (No Tools)92.9023 / 274Max (No Tools)-2

General Knowledge

DeepSeek-V4.1-Flash 1/1
BenchmarkDeepSeek-V4.1-FlashKimi K3Diff
HLE63.903 / 197Max (With Tools)59.808 / 197Max (With Tools)+4.10

Math and Reasoning

Even 1/1
BenchmarkDeepSeek-V4.1-FlashKimi K3Diff
MathArena Apex65.601 / 3Max (No Tools)65.601 / 3Max (No Tools)

Productivity Knowledge

DeepSeek-V4.1-Flash 1/1
BenchmarkDeepSeek-V4.1-FlashKimi K3Diff
AutomationBench54.801 / 17Max (With Tools)46.707 / 17Max (With Tools)+8.10

Specs

FieldDeepSeek-V4.1-FlashKimi K3
PublisherDeepSeek-AIMoonshot AI
Release date2026-09-102026-07-16
Model typeMultimodal modelReasoning model
ArchitectureMoEMoE
Parameters552B2.8T
Context length1M1M
Max output384K1M

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemDeepSeek-V4.1-FlashKimi K3
Text input¥1 / 1M tokens¥20 / 1M tokens
Text output¥4 / 1M tokens¥100 / 1M tokens
Cache read¥0.02 / 1M tokens¥2 / 1M tokens

Summary

  • DeepSeek-V4.1-Flashleads in:AI Agent - Tool Usage (4/4), Coding and Software Engineer (3/3), Multimodal Understanding (3/3), Agent Level Benchmark (1/1), General Knowledge (1/1), Productivity Knowledge (1/1)
  • Kimi K3leads in:General Evaluation (1/1)
  • Tied in:Math and Reasoning

On average across the 15 shared benchmarks, DeepSeek-V4.1-Flash scores 6.35 higher.

Largest single-benchmark gap: Terminal-Bench 4.0 — DeepSeek-V4.1-Flash 31.20 vs Kimi K3 12.60 (+18.60).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.