DeepSeek-V4.1-FlashvsDeepSeek-V4-Flash
Across 12 shared benchmarks, DeepSeek-V4.1-Flash leads overall: DeepSeek-V4.1-Flash wins 12, DeepSeek-V4-Flash wins 0, with 0 ties and an average score difference of +29.58.
DeepSeek-V4.1-Flash
DeepSeek-AI · 2026-09-10 · Multimodal model
DeepSeek-V4-Flash
DeepSeek-AI · 2026-04-24 · Reasoning model
DeepSeek-V4.1-Flash12 wins(100%)(0%)0 winsDeepSeek-V4-Flash
Benchmark scores
Grouped by capability, sorted by largest gap within each. 12 shared benchmarks.
AI Agent - Tool Usage
DeepSeek-V4.1-Flash 4/4| Benchmark | DeepSeek-V4.1-Flash | DeepSeek-V4-Flash | Diff |
|---|---|---|---|
| Terminal-Bench 4.0 | 31.207 / 16Max (With Tools) | 716 / 16Max (With Tools) | +24.20 |
| Terminal-Bench 3.0 | 304 / 11Max (With Tools) | 7.6011 / 11Max (With Tools) | +22.40 |
| CyberGym | 88.101 / 8Max (With Tools) | 76.707 / 8Max (With Tools) | +11.40 |
| Terminal-Bench 2.1 | 90.601 / 53Max (With Tools) | 82.7024 / 53Max (With Tools) | +7.90 |
Coding and Software Engineer
DeepSeek-V4.1-Flash 4/4| Benchmark | DeepSeek-V4.1-Flash | DeepSeek-V4-Flash | Diff |
|---|---|---|---|
| CodeForces | 3,4711 / 21Max (No Tools) | 3,2893 / 21Max (No Tools) | +182 |
| SEC-Bench Pro | 62.804 / 5Max (With Tools) | 30.905 / 5Max (With Tools) | +31.90 |
| DeepSWE | 74.202 / 38Max (With Tools) | 54.4026 / 38Max (With Tools) | +19.80 |
| NL2Repo-Bench | 65.402 / 16Max (With Tools) | 54.2012 / 16Max (With Tools) | +11.20 |
Agent Capability
DeepSeek-V4.1-Flash 1/1| Benchmark | DeepSeek-V4.1-Flash | DeepSeek-V4-Flash | Diff |
|---|---|---|---|
| ExploitGym (budget unspecified) | 15.302 / 4Max (With Tools) | 1.804 / 4Max (With Tools) | +13.50 |
Agent Level Benchmark
DeepSeek-V4.1-Flash 1/1| Benchmark | DeepSeek-V4.1-Flash | DeepSeek-V4-Flash | Diff |
|---|---|---|---|
| Agents' Last Exam | 31.806 / 19Max (With Tools) | 25.2016 / 19Max (With Tools) | +6.60 |
Math and Reasoning
DeepSeek-V4.1-Flash 1/1| Benchmark | DeepSeek-V4.1-Flash | DeepSeek-V4-Flash | Diff |
|---|---|---|---|
| MathArena Apex | 65.601 / 3Max (No Tools) | 58.603 / 3Max (No Tools) | +7 |
Productivity Knowledge
DeepSeek-V4.1-Flash 1/1| Benchmark | DeepSeek-V4.1-Flash | DeepSeek-V4-Flash | Diff |
|---|---|---|---|
| AutomationBench | 54.801 / 17Max (With Tools) | 37.7010 / 17Max (With Tools) | +17.10 |
Specs
| Field | DeepSeek-V4.1-Flash | DeepSeek-V4-Flash |
|---|---|---|
| Publisher | DeepSeek-AI | DeepSeek-AI |
| Release date | 2026-09-10 | 2026-04-24 |
| Model type | Multimodal model | Reasoning model |
| Architecture | MoE | MoE |
| Parameters | 552B | 284B |
| Context length | 1M | 1M |
| Max output | 384K | 384K |
API pricing
Prices use DataLearner records when available; missing fields are not inferred.
| Item | DeepSeek-V4.1-Flash | DeepSeek-V4-Flash |
|---|---|---|
| Text input | ¥1 / 1M tokens | $0.14 / 1M tokens |
| Text output | ¥4 / 1M tokens | $0.28 / 1M tokens |
| Cache read | ¥0.02 / 1M tokens | $0.0028 / 1M tokens |
Summary
- DeepSeek-V4.1-Flashleads in:AI Agent - Tool Usage (4/4), Coding and Software Engineer (4/4), Agent Capability (1/1), Agent Level Benchmark (1/1), Math and Reasoning (1/1), Productivity Knowledge (1/1)
On average across the 12 shared benchmarks, DeepSeek-V4.1-Flash scores 29.58 higher.
Largest single-benchmark gap: CodeForces — DeepSeek-V4.1-Flash 3,471 vs DeepSeek-V4-Flash 3,289 (+182).
Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.