DataLearner logo

DeepSeek-V4-Flash Benchmark Details

DeepSeek-V4-Flash currently shows benchmark results led by LiveCodeBench (6 / 250, score 91.60), IF Bench (11 / 282, score 79.20), HLE (43 / 563, score 51.50). This page also compares it with 3 competitor models and 2 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

DeepSeek-V4-Flash

Benchmark Results

Thinking
Tool usage
Internet

General Knowledge

16 evaluations
Benchmark / mode
Score
Rank/total
MMLU Pro
Standard Mode
83
47 / 176
86.40
18 / 176
86.20
19 / 176
LiveBench
Standard Mode
65.48
52 / 117
HLE
Standard Mode
8.10
418 / 563
HLE
Standard Mode
7.80
424 / 563
HLE
High
30.30
200 / 563
HLE
High
29.40
210 / 563
HLE
HighTools
40.30
130 / 563
HLE
Max
34.80
171 / 563
HLE
Max
34.80
171 / 563
HLE
MaxTools
51.50
43 / 563
HLE
Thinking Level · Extra HighTools
45.10
82 / 563
CritPt
Standard Mode
0.30
182 / 200
CritPt
High
3.40
109 / 200
7.10
87 / 200

Other

3 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
71.20
311 / 462
86.70
136 / 462
89.40
96 / 462

Coding and Software Engineer

21 evaluations
Benchmark / mode
Score
Rank/total
2816
6 / 21
3289
3 / 21
1576.54
6 / 35
LiveCodeBench
Standard Mode
55.20
151 / 250
88.40
16 / 250
91.60
6 / 250
SWE-bench Verified
Standard ModeTools
73.70
45 / 116
78.60
23 / 116
SWE-bench Verified
Thinking Level · Extra HighTools
79
20 / 116
SWE-bench Multilingual
Standard ModeTools
69.70
24 / 30
70.20
22 / 30
SWE-bench Multilingual
Thinking Level · Extra HighTools
73.30
17 / 30
54.20
12 / 16
DeepSWE
MaxTools
53.32
56 / 85
SWE-Bench Pro - Public
Standard ModeTools
49.10
50 / 62
52.30
42 / 62
SWE-Bench Pro - Public
Thinking Level · Extra HighTools
52.60
40 / 62
40.20
101 / 130
45.30
89 / 130
WeirdML v2
HighTools
43.76
44 / 52
30.90
6 / 6

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1555.70
46 / 106

Common Sense Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Standard Mode
61.10
29 / 93
SimpleBench
Standard Mode
46.30
57 / 93

Agent Level Benchmark

9 evaluations
Benchmark / mode
Score
Rank/total
τ²-Bench - Telecom
Standard ModeTools
94.40
35 / 264
95.60
26 / 264
95
32 / 264
Terminal Bench Hard
Standard ModeTools
34.10
82 / 244
38.60
57 / 244
35.60
70 / 244
τ³-Banking
HighTools
26.20
71 / 164
τ³-Banking
MaxTools
30.90
57 / 164
25.20
16 / 19

Instruction Following

3 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Standard Mode
47.20
170 / 282
73.50
43 / 282
79.20
11 / 282

AI Agent - Information Search

2 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
HighTools
53.50
42 / 57
BrowseComp
Thinking Level · Extra HighTools
73.20
30 / 57

AI Agent - Tool Usage

10 evaluations
Benchmark / mode
Score
Rank/total
CyberGym
MaxTools
76.70
7 / 8
70.30
11 / 11
56.90
117 / 192
61.80
105 / 192
Terminal Bench 2.0
Standard ModeTools
49.10
36 / 48
56.60
28 / 48
Terminal Bench 2.0
Thinking Level · Extra HighTools
56.90
26 / 48
7.60
11 / 11
3
60 / 86
2.50
63 / 86

Text Embedding

2 evaluations
Benchmark / mode
Score
Rank/total
Context Arena
Standard Mode
26.47
117 / 126
Context Arena
Thinking Enabled
69.42
66 / 126

Math and Reasoning

7 evaluations
Benchmark / mode
Score
Rank/total
AIME 2026
MaxTools
95.83
10 / 29
93.94
6 / 10
IMO-AnswerBench
Standard Mode
41.90
23 / 24
85.10
12 / 24
88.40
7 / 24
58.60
7 / 17
27.08
12 / 17

Long Context

3 evaluations
Benchmark / mode
Score
Rank/total
AA-LCR
Standard Mode
41.70
156 / 170
AA-LCR
High
72
95 / 170
74.30
86 / 170

Productivity Knowledge

7 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA v2
HighTools
1083
84 / 105
GDPval-AA v2
MaxTools
1116
79 / 105
AA-Briefcase
HighTools
1052
57 / 83
AA-Briefcase
MaxTools
833
75 / 83
81.33
31 / 43
37.70
10 / 17
AA-AnalystAgent
MaxToolsInternet
25
16 / 29

Multimodal Understanding

2 evaluations
Benchmark / mode
Score
Rank/total
10.40
78 / 118
10.80
77 / 118

Other

1 evaluations
Benchmark / mode
Score
Rank/total

Claw-style Agent Evaluation

1 evaluations
Benchmark / mode
Score
Rank/total
81.74
8 / 45

Competitor Comparison

Benchmark scores for DeepSeek-V4-Flash compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkDeepSeek-V4-FlashCurrentGLM 5.1Qwen3.6-27BGemini 3.5 Flash
CritPt
Score
综合评估
7.10Thinking Level · High
4.60Thinking Enabled
1.10Thinking Enabled
13.10Thinking Level · High
HLE
Accuracy
综合评估
51.50Thinking Level · High | Tools
52.30Thinking Enabled | Tools
24.00Thinking Enabled
42.70Thinking Level · High
LiveBench
Accuracy
综合评估
65.48Standard Mode
70.18Standard Mode
64.03Standard Mode
74.64Thinking Level · High
MMLU Pro
Accuracy
综合评估
86.40Thinking Level · High
--
86.20Thinking Enabled
--
GPQA Diamond
Accuracy
科学与综合推理
89.40Thinking Level · High
86.20Thinking Enabled
87.80Thinking Enabled
92.80Thinking Level · High
DeepSWE
Pass@1 (DeepSWE v1.1)
编程与软件工程
53.32Thinking Level · High | Tools
--
--
37.00Thinking Level · Medium | Tools
LiveCodeBench
Pass @K
编程与软件工程
91.60Thinking Level · High
--
83.90Thinking Enabled
--
SciCode
Score
编程与软件工程
45.30Thinking Level · High
44.80Thinking Enabled
42.80Thinking Enabled
53.90Thinking Level · High
SWE-bench Multilingual
Accuracy
编程与软件工程
73.30Thinking Level · Extra High | Tools
--
71.30Thinking Enabled | Tools
--
SWE-Bench Pro - Public
Accuracy
编程与软件工程
52.60Thinking Level · Extra High | Tools
58.40Thinking Enabled | Tools
53.50Thinking Enabled | Tools
55.10Thinking Level · High | Tools
SWE-bench Verified
Accuracy
编程与软件工程
79.00Thinking Level · Extra High | Tools
--
77.20Thinking Enabled | Tools
--
Text Arena (Coding)
Arena Score
编程与软件工程
1576.54Thinking Level · High
1534.00Standard Mode
--
--
21 additional benchmarks remain in the chart above.

Standard API Pricing: DeepSeek-V4-Flash vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
DeepSeek-V4-Flash
DeepSeek-AI$0.14 / 1M tokens$0.28 / 1M tokens
GLM 5.1
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens
Gemini 3.5 Flash
DeepMind$1.5 / 1M tokens$9 / 1M tokens

Version History

How each version of the DeepSeek-V4-Flash series stacks up on benchmark tests

DeepSeek-V4-FlashDeepSeek V4DeepSeek V3.2
Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkDeepSeek-V4-FlashCurrentDeepSeek V3.2
CritPt
Score
综合评估
7.10Thinking Level · High
2.90Thinking Enabled
HLE
Accuracy
综合评估
51.50Thinking Level · High | Tools
25.10Thinking Enabled
LiveBench
Accuracy
综合评估
65.48Standard Mode
62.20Thinking Enabled
GPQA Diamond
Accuracy
科学与综合推理
89.40Thinking Level · High
82.40Thinking Enabled
CodeForces
Accuracy
编程与软件工程
3289.00Thinking Level · High
2386.00Thinking Enabled
LiveCodeBench
Pass @K
编程与软件工程
91.60Thinking Level · High
83.30Thinking Enabled
SWE-Bench Pro - Public
Accuracy
编程与软件工程
52.60Thinking Level · Extra High | Tools
40.90Thinking Enabled
SWE-bench Verified
Accuracy
编程与软件工程
79.00Thinking Level · Extra High | Tools
73.10Thinking Enabled | Tools
Creative Writing
Elo、大模型评判两两对战
写作和创作
1555.70Standard Mode
1511.20Standard Mode
Terminal Bench Hard
Accuracy
Agent能力评测
38.60Thinking Level · High | Tools
35.60Thinking Enabled | Tools
τ²-Bench - Telecom
Accuracy
Agent能力评测
95.60Thinking Level · High | Tools
90.60Thinking Enabled | Tools
IF Bench
Accuracy
指令跟随
79.20Thinking Level · High
60.70Thinking Enabled
5 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: CritPt · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the DeepSeek-V4-Flash Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
DeepSeek-V4-Flash
DeepSeek-AI$0.14 / 1M tokens$0.28 / 1M tokens
DeepSeek V3.2
DeepSeek-AI$0.28 / 1M tokens$0.42 / 1M tokens