DataLearner logo

DeepSeek-V4.1-Flash Benchmark Details

DeepSeek-V4.1-Flash currently shows benchmark results led by HLE (3 / 197, score 63.90), Terminal-Bench 2.1 (1 / 53, score 90.60), CodeForces (1 / 21, score 3471). This page also compares it with 3 competitor models and 2 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

DeepSeek-V4.1-Flash

Benchmark Results

Thinking
Tool usage

General Knowledge

3 evaluations
Benchmark / mode
Score
Rank/total
HLE
Max
39.10
78 / 197
HLE
Max
36.80
88 / 197
HLE
MaxTools
63.90
3 / 197

Other

1 evaluations
Benchmark / mode
Score
Rank/total
90.90
39 / 274

Coding and Software Engineer

5 evaluations
Benchmark / mode
Score
Rank/total
3471
1 / 21
DeepSWE
MaxTools
74.20
2 / 38
65.40
2 / 16
62.80
4 / 5
20.30
8 / 11

AI Agent - Tool Usage

4 evaluations
Benchmark / mode
Score
Rank/total
90.60
1 / 53
CyberGym
MaxTools
88.10
1 / 8
31.20
7 / 16
30
4 / 11

Agent Level Benchmark

1 evaluations
Benchmark / mode
Score
Rank/total
31.80
6 / 19

Productivity Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
54.80
1 / 17

Multimodal Understanding

3 evaluations
Benchmark / mode
Score
Rank/total
BabyVision
MaxTools
89.60
3 / 9
Chartography
MaxTools
78.90
4 / 8
49
3 / 6

Math and Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
65.60
1 / 3

Other

1 evaluations
Benchmark / mode
Score
Rank/total

Competitor Comparison

Benchmark scores for DeepSeek-V4.1-Flash compared against top models in its class

DeepSeek-V4.1-FlashKimi K3GLM-5.3Claude Opus 5
Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkDeepSeek-V4.1-FlashCurrentKimi K3GLM-5.3Claude Opus 5
HLE
Accuracy
综合评估
63.90Thinking Level · High | Tools
59.80Thinking Level · High | Tools
62.50Thinking Level · High | Tools
63.60Thinking Level · High | Tools
GPQA Diamond
Accuracy
科学与综合推理
90.90Thinking Level · High
92.90Thinking Level · High
88.10Thinking Level · High
93.40Thinking Level · High
DeepSWE
Pass@1 (DeepSWE v1.1)
编程与软件工程
74.20Thinking Level · High | Tools
67.50Thinking Level · High | Tools
66.90Thinking Level · High | Tools
74.00Thinking Level · High | Tools
NL2Repo-Bench
Average test pass rate
编程与软件工程
65.40Thinking Level · High | Tools
58.00Thinking Level · High | Tools
58.00Thinking Level · High | Tools
75.30Thinking Level · High | Tools
Program Bench
Score
编程与软件工程
20.30Thinking Level · High | Tools
17.50Thinking Level · High | Tools
19.00Thinking Level · High | Tools
37.00Thinking Level · High | Tools
CyberGym
Score
AI Agent - 工具使用
88.10Thinking Level · High | Tools
80.00Thinking Level · High | Tools
84.50Thinking Level · High | Tools
--
Terminal-Bench 2.1
Accuracy
AI Agent - 工具使用
90.60Thinking Level · High | Tools
88.30Thinking Level · High | Tools
88.20Thinking Level · High | Tools
89.10Thinking Level · High | Tools
Terminal-Bench 3.0
Accuracy
AI Agent - 工具使用
30.00Thinking Level · High | Tools
17.70Thinking Level · High | Tools
28.30Thinking Level · High | Tools
43.30Thinking Level · High | Tools
Terminal-Bench 4.0
Resolution rate (%)
AI Agent - 工具使用
31.20Thinking Level · High | Tools
12.60Thinking Level · High | Tools
37.90Thinking Level · High | Tools
51.80Thinking Level · High | Tools
Agents' Last Exam
Score
Agent能力评测
31.80Thinking Level · High | Tools
27.60Thinking Level · High | Tools
28.50Thinking Level · High | Tools
28.60Thinking Level · High | Tools
AutomationBench
Pass Rate
生产力知识
54.80Thinking Level · High | Tools
46.70Thinking Level · High | Tools
48.80Thinking Level · High | Tools
50.30Thinking Level · High | Tools
BabyVision
Accuracy
多模态理解
89.60Thinking Level · High | Tools
85.70Thinking Level · High | Tools
--
94.10Thinking Level · High | Tools
4 additional benchmarks remain in the chart above.

Standard API Pricing: DeepSeek-V4.1-Flash vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

Kimi K3
Supplier: Moonshot AI
Standard input: ¥20 / 1M tokens
Standard output: ¥100 / 1M tokens
GLM-5.3
Supplier: 智谱AI
Standard input: $1.4 / 1M tokens
Standard output: $4.4 / 1M tokens
Claude Opus 5
Supplier: Anthropic
Standard input: $5 / 1M tokens
Standard output: $25 / 1M tokens
ModelSupplierStandard inputStandard outputBase price applies to
Kimi K3
Moonshot AI¥20 / 1M tokens¥100 / 1M tokens
GLM-5.3
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens
Claude Opus 5
Anthropic$5 / 1M tokens$25 / 1M tokens

Version History

How each version of the DeepSeek-V4.1-Flash series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkDeepSeek-V4.1-FlashCurrentDeepSeek-V4-Flash-Vision-ExpDeepSeek-V4-Flash
HLE
Accuracy
综合评估
63.90Thinking Level · High | Tools
--
51.50Thinking Level · High | Tools
GPQA Diamond
Accuracy
科学与综合推理
90.90Thinking Level · High
--
89.90Thinking Level · High
CodeForces
Accuracy
编程与软件工程
3471.00Thinking Level · High
--
3289.00Thinking Level · High
DeepSWE
Pass@1 (DeepSWE v1.1)
编程与软件工程
74.20Thinking Level · High | Tools
59.30Thinking Level · High | Tools
54.40Thinking Level · High | Tools
NL2Repo-Bench
Average test pass rate
编程与软件工程
65.40Thinking Level · High | Tools
57.70Thinking Level · High | Tools
54.20Thinking Level · High | Tools
SEC-Bench Pro
Score (%)
编程与软件工程
62.80Thinking Level · High | Tools
--
30.90Thinking Level · High | Tools
CyberGym
Score
AI Agent - 工具使用
88.10Thinking Level · High | Tools
75.30Thinking Level · High | Tools
76.70Thinking Level · High | Tools
Terminal-Bench 2.1
Accuracy
AI Agent - 工具使用
90.60Thinking Level · High | Tools
83.90Thinking Level · High | Tools
82.70Thinking Level · High | Tools
Terminal-Bench 3.0
Accuracy
AI Agent - 工具使用
30.00Thinking Level · High | Tools
--
7.60Thinking Level · High | Tools
Terminal-Bench 4.0
Resolution rate (%)
AI Agent - 工具使用
31.20Thinking Level · High | Tools
--
7.00Thinking Level · High | Tools
Agents' Last Exam
Score
Agent能力评测
31.80Thinking Level · High | Tools
27.30Thinking Level · High | Tools
25.20Thinking Level · High | Tools
AutomationBench
Pass Rate
生产力知识
54.80Thinking Level · High | Tools
25.70Thinking Level · High | Tools
37.70Thinking Level · High | Tools
4 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: HLE · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the DeepSeek-V4.1-Flash Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
DeepSeek-V4-Flash
DeepSeek-AI$0.14 / 1M tokens$0.28 / 1M tokens