DataLearner logo
DE

DeepSeek-V4-Flash

PreviewReasoning modelCoding modelDeepSeek FlashDeepSeek V4

DeepSeek V4 Flash

Release date: 2026-04-24Updated: 2026-09-10Views: 10,129
Parameters
284B
Context length
1M
Multilingual
Supported
Reasoning ability
5/5

V4 Flash remains a distinct 284B/13B MoE with its existing lifecycle and original weights. Updated 0731 max-effort scores come from the September 10 release table. The deepseek-v4-flash API alias temporarily routes to V4.1 Flash; this does not change the original model or weights. USD price rows on this page are historical, not current quotes for routed requests. Current routed CNY rates per million tokens: off-peak cache-hit/cache-miss/output 0.02/1/4; peak 0.04/2/8 (weekdays 09:00-12:00, 14:00-18:00 Asia/Shanghai).

Data sourced primarily from official releases (GitHub, Hugging Face, papers), then benchmark leaderboards, then third-party evaluators. Learn about our data methodology

DeepSeek-V4-Flash

Model basics

Reasoning traces
Supported
Thinking modes
Thinking Level · Max (Default)Standard ModeThinking Level · High
Context length
1M tokens
Max output length
384K tokens
Model type
Reasoning model
Modality (in / out)
Text → Text
Release date
2026-04-24
Model file size
No data
MoE architecture
Yes
Total params / Active params
284B / 13B
Knowledge cutoff
No data
DeepSeek-V4-Flash

Open source & experience

Code license
Weights license
MIT License- Commercial use permitted
GitHub repo
N/A
DeepSeek-V4-Flash

Official resources

DeepSeek-V4-Flash

API details

API speed
4/5
💡Default unit: $/1M tokens. If vendors use other units, follow their published pricing.
Standard
TypeConditionInputOutput
Text-$0.140/ 1M tokens$0.280/ 1M tokens
Cache PricingPrompt Cache
TypeTTLWriteRead
Text-$0.0028/ 1M tokens

“—” means the modality is not billed in that direction, or the vendor has not published a price for it.

DeepSeek-V4-Flash

Benchmark Results

DeepSeek-V4-Flash currently shows benchmark results led by LiveCodeBench (6 / 250, score 91.60), IF Bench (11 / 282, score 79.20), HLE (43 / 563, score 51.50). This page also consolidates core specs, context limits, and API pricing so you can evaluate the model from benchmark results and deployment constraints together.

Thinking
Tool usage
Internet

General Knowledge

16 evaluations
Benchmark / mode
Score
Rank/total
MMLU Pro
Standard Mode
83
47 / 176
86.40
18 / 176
86.20
19 / 176
LiveBench
Standard Mode
65.48
52 / 117
HLE
Standard Mode
8.10
418 / 563
HLE
Standard Mode
7.80
424 / 563
HLE
High
30.30
200 / 563
HLE
High
29.40
210 / 563
HLE
HighTools
40.30
130 / 563
HLE
Max
34.80
171 / 563
HLE
Max
34.80
171 / 563
HLE
MaxTools
51.50
43 / 563
HLE
Extra-HighTools
45.10
82 / 563
CritPt
Standard Mode
0.30
182 / 200
CritPt
High
3.40
109 / 200
7.10
87 / 200

General Evaluation

3 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
71.20
311 / 462
86.70
136 / 462
89.40
96 / 462

Coding and Software Engineer

21 evaluations
Benchmark / mode
Score
Rank/total
2816
6 / 21
3289
3 / 21
1576.54
6 / 35
LiveCodeBench
Standard Mode
55.20
151 / 250
88.40
16 / 250
91.60
6 / 250
SWE-bench Verified
Standard ModeTools
73.70
45 / 116
78.60
23 / 116
SWE-bench Verified
Extra-HighTools
79
20 / 116
SWE-bench Multilingual
Standard ModeTools
69.70
24 / 30
70.20
22 / 30
SWE-bench Multilingual
Extra-HighTools
73.30
17 / 30
54.20
12 / 16
DeepSWE
MaxTools
53.32
56 / 85
SWE-Bench Pro - Public
Standard ModeTools
49.10
50 / 62
52.30
42 / 62
SWE-Bench Pro - Public
Extra-HighTools
52.60
40 / 62
40.20
101 / 130
45.30
89 / 130
WeirdML v2
HighTools
43.76
44 / 52
30.90
6 / 6

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1555.70
46 / 106

Common Sense Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Standard Mode
61.10
29 / 93
SimpleBench
Standard Mode
46.30
57 / 93

Agent Level Benchmark

9 evaluations
Benchmark / mode
Score
Rank/total
τ²-Bench - Telecom
Standard ModeTools
94.40
35 / 264
95.60
26 / 264
95
32 / 264
Terminal Bench Hard
Standard ModeTools
34.10
82 / 244
38.60
57 / 244
35.60
70 / 244
τ³-Banking
HighTools
26.20
71 / 164
τ³-Banking
MaxTools
30.90
57 / 164
25.20
16 / 19

Instruction Following

3 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Standard Mode
47.20
170 / 282
73.50
43 / 282
79.20
11 / 282

AI Agent - Information Search

2 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
HighTools
53.50
42 / 57
BrowseComp
Extra-HighTools
73.20
30 / 57

AI Agent - Tool Usage

10 evaluations
Benchmark / mode
Score
Rank/total
CyberGym
MaxTools
76.70
7 / 8
70.30
11 / 11
56.90
117 / 192
61.80
105 / 192
Terminal Bench 2.0
Standard ModeTools
49.10
36 / 48
56.60
28 / 48
Terminal Bench 2.0
Extra-HighTools
56.90
26 / 48
7.60
11 / 11
3
60 / 86
2.50
63 / 86

Text Embedding

2 evaluations
Benchmark / mode
Score
Rank/total
Context Arena
Standard Mode
26.47
117 / 126
Context Arena
Thinking Mode
69.42
66 / 126

Math and Reasoning

7 evaluations
Benchmark / mode
Score
Rank/total
AIME 2026
MaxTools
95.83
10 / 29
93.94
6 / 10
IMO-AnswerBench
Standard Mode
41.90
23 / 24
85.10
12 / 24
88.40
7 / 24
58.60
7 / 17
27.08
12 / 17

Long Context

3 evaluations
Benchmark / mode
Score
Rank/total
AA-LCR
Standard Mode
41.70
156 / 170
AA-LCR
High
72
95 / 170
74.30
86 / 170

Productivity Knowledge

7 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA v2
HighTools
1083
84 / 105
GDPval-AA v2
MaxTools
1116
79 / 105
AA-Briefcase
HighTools
1052
57 / 83
AA-Briefcase
MaxTools
833
75 / 83
81.33
31 / 43
37.70
10 / 17
AA-AnalystAgent
MaxToolsInternet
25
16 / 29

Multimodal Understanding

2 evaluations
Benchmark / mode
Score
Rank/total
10.40
78 / 118
10.80
77 / 118

Agent Capability

1 evaluations
Benchmark / mode
Score
Rank/total

Claw-style Agent Evaluation

1 evaluations
Benchmark / mode
Score
Rank/total
81.74
8 / 45

Compare with other models

DeepSeek-V4-Flash

Publisher

DeepSeek V4 Flash

Model Overview

September 10, 2026 update: V4 Flash 0731

This page preserves DeepSeek V4 Flash as a distinct model with its existing lifecycle status. No retirement, deletion, or model-page redirect is applied. The V4 Flash 0731 column in the supplied September 10 release table updates its comparison results; it must not be confused with V4.1 Flash.

Updated max-effort results

GPQA Diamond 89.9; HLE pure-text subset 37.8; Codeforces rating 3289; MathArena Apex 58.6; Terminal-Bench 2.1/3.0/4.0: 82.7/7.6/7.0; DeepSWE v1.1 54.4; NL2Repo-Bench 54.2; CyberGym 76.7; SEC-Bench Pro 30.9; ExploitGym 1.8 (budget unspecified); HLE with tools 51.5; Automation-Bench 37.7; Agents’ Last Exam 25.2.

The supplied clarification identifies max effort, with tool use tracked separately and internet access not inferred. Pure-text HLE is kept separate from full HLE, and budget-unspecified ExploitGym is separate from 2h/6h/unlimited protocols. Dashes for ProgramBench and visual evaluations mean no result supplied, not zero. The original 0731 Toolathlon Verified score of 70.3 and other historical modes and third-party results are retained. The new Automation-Bench value 37.7 replaces 25.1 in the max-with-tools entry; the old release value remains documented in the local update record.

Architecture and open-weight identity

Existing specifications remain: 284B total and 13B active parameters, MoE, text-only input and output, 1M context and 384K maximum output. The V4 architecture combines Compressed Sparse Attention, Heavily Compressed Attention, mHC connections and the Muon optimizer. The original V4 Flash weights and their existing MIT license remain distinct from the new API route. The 0731 release updated API post-training, not the previously published weights.

Compatibility routing and historical prices

According to the supplied September 10 materials, deepseek-v4-flash temporarily routes to V4.1 Flash; the recommended API name is deepseek-flash. Calls through the compatibility name execute the new model and use its peak/off-peak prices. This does not alter the original model’s lifecycle, weights, or historical benchmarks.

The price rows retained for the original model are historical: USD 0.0028 cache-hit input, 0.14 cache-miss input, and 0.28 output per million tokens. They are not current quotes for requests routed through the compatibility alias. V4.1 Flash separately lists off-peak CNY 0.02/1/4 and peak CNY 0.04/2/8 for those three categories. Peak hours are Monday–Friday 09:00–12:00 and 14:00–18:00 Asia/Shanghai.

Historical reasoning and API support

Original off/high/max thinking modes and the existing default are preserved. Original API support includes OpenAI/Anthropic compatibility, Responses API, JSON Output, Tool Calls, prefix completion, and non-thinking FIM. V4 Flash remains a text model; routing its old API name to a vision model does not give its own weights image support.

Source and evaluation notes

Updated scores and routing information come from the supplied DeepSeek September 10, 2026 release materials without independent source verification in this update. Original 0731 public Code Agent notes specified DeepSeek Harness minimal mode, max effort, temperature=1.0 and top_p=0.95; this does not establish identical settings for every historical benchmark. Reference: DeepSeek change log.

DeepSeek-V4-Flash

FAQ

Was the V4 Flash model record retired or removed?

No. This update preserves its lifecycle, independent record, and original weights while updating information and scores.

What does the deepseek-v4-flash API name call now?

The supplied materials say the compatibility name temporarily routes to V4.1 Flash at its prices. Original V4 Flash USD price rows are historical.

DataLearner on WeChat

Follow DataLearner on WeChat for AI model updates and research notes.

DataLearner WeChat QR code