DataLearner logo

DeepSeek-V4-Pro Benchmark Details

DeepSeek-V4-Pro currently shows benchmark results led by LiveCodeBench (1 / 128, score 93.50), MMLU Pro (11 / 134, score 87.50), SWE-bench Verified (11 / 116, score 80.60). This page also compares it with 3 competitor models and 4 predecessor or same-series models, including performance and pricing views when available. 1 source link is attached for reference.

Benchmark Results

DeepSeek-V4-Pro

Benchmark Results

Thinking
Tool usage

General Knowledge

9 evaluations
Benchmark / mode
Score
Rank/total
MMLU Pro
Standard Mode
82.90
49 / 134
87.10
13 / 134
87.50
11 / 134
LiveBench
Standard Mode
71.57
32 / 117
HLE
Standard Mode
7.70
181 / 197
HLE
High
34.50
95 / 197
HLE
HighTools
44.70
53 / 197
HLE
Max
37.70
82 / 197
HLE
Thinking Level · Extra HighTools
48.20
45 / 197

Other

3 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
72.90
185 / 274
89.10
59 / 274
90.10
50 / 274

Coding and Software Engineer

16 evaluations
Benchmark / mode
Score
Rank/total
2919
5 / 21
3206
4 / 21
LiveCodeBench
Standard Mode
56.80
81 / 128
89.80
8 / 128
93.50
1 / 128
SWE-bench Verified
Standard ModeTools
73.60
46 / 116
79.40
19 / 116
SWE-bench Verified
Thinking Level · Extra HighTools
80.60
11 / 116
SWE-bench Multilingual
Standard ModeTools
69.80
22 / 29
74.10
12 / 29
SWE-bench Multilingual
Thinking Level · Extra HighTools
76.20
10 / 29
DeepSWE
Thinking Level · Extra HighTools
62.70
20 / 38
NL2Repo-Bench
Thinking Level · Extra HighTools
61.50
4 / 16
SWE-Bench Pro - Public
Standard ModeTools
52.10
43 / 62
54.40
34 / 62
SWE-Bench Pro - Public
Thinking Level · Extra HighTools
55.40
31 / 62

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1552.10
47 / 106

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Standard Mode
50.90
50 / 94

AI Agent - Information Search

2 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
HighTools
80.40
18 / 57
BrowseComp
Thinking Level · Extra HighTools
83.40
15 / 57

AI Agent - Tool Usage

6 evaluations
Benchmark / mode
Score
Rank/total
Terminal-Bench 2.1
Thinking Level · Extra HighTools
87.90
11 / 53
CyberGym
Thinking Level · Extra HighTools
83.30
4 / 8
Toolathlon-Verified
Thinking Level · Extra HighTools
74.10
5 / 11
Terminal Bench 2.0
Standard ModeTools
59.10
22 / 48
63.30
14 / 48
Terminal Bench 2.0
Thinking Level · Extra HighTools
67.90
9 / 48

Text Embedding

2 evaluations
Benchmark / mode
Score
Rank/total
Context Arena
Standard Mode
31.43
115 / 126
Context Arena
Thinking Enabled
76.09
48 / 126

Math and Reasoning

5 evaluations
Benchmark / mode
Score
Rank/total
IMO-AnswerBench
Standard Mode
35.30
24 / 24
88
8 / 24
89.80
5 / 24
45.26
31 / 58

Productivity Knowledge

3 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA
Thinking Level · Extra HighTools
1554
4 / 21
AA-Briefcase
MaxTools
1286.37
13 / 20
AutomationBench
Thinking Level · Extra HighTools
31.80
12 / 17

Agent Level Benchmark

1 evaluations
Benchmark / mode
Score
Rank/total
Agents' Last Exam
Thinking Level · Extra HighTools
25.70
15 / 19

Competitor Comparison

Benchmark scores for DeepSeek-V4-Pro compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkDeepSeek-V4-ProCurrentGLM 5.1Kimi K2.6Qwen3.8-Max
HLE
Accuracy
综合评估
48.20Thinking Level · Extra High | Tools
52.30Thinking Enabled | Tools
54.00Thinking Enabled | Tools
56.20Thinking Level · Extra High | Tools
LiveBench
Accuracy
综合评估
71.57Standard Mode
70.18Standard Mode
70.54Thinking Enabled
--
GPQA Diamond
Accuracy
科学与综合推理
90.10Thinking Level · High
86.20Thinking Enabled
90.50Thinking Enabled
92.60Thinking Level · Extra High
DeepSWE
Pass@1 (DeepSWE v1.1)
编程与软件工程
62.70Thinking Level · Extra High | Tools
--
--
56.60Thinking Level · Extra High | Tools
LiveCodeBench
Pass @K
编程与软件工程
93.50Thinking Level · High
--
89.60Thinking Enabled
--
NL2Repo-Bench
Average test pass rate
编程与软件工程
61.50Thinking Level · Extra High | Tools
--
--
55.90Thinking Level · Extra High | Tools
SWE-bench Multilingual
Accuracy
编程与软件工程
76.20Thinking Level · Extra High | Tools
--
76.70Thinking Enabled | Tools
--
SWE-Bench Pro - Public
Accuracy
编程与软件工程
55.40Thinking Level · Extra High | Tools
58.40Thinking Enabled | Tools
58.60Thinking Enabled | Tools
67.70Thinking Level · Extra High | Tools
SWE-bench Verified
Accuracy
编程与软件工程
80.60Thinking Level · Extra High | Tools
--
80.20Thinking Enabled | Tools
--
Creative Writing
Elo、大模型评判两两对战
写作和创作
1552.10Standard Mode
1589.20Standard Mode
1721.30Standard Mode
1840.90Standard Mode
SimpleBench
Score (AVG@5)
常识推理
50.90Standard Mode
55.10Standard Mode
--
62.50Thinking Level · Extra High
BrowseComp
Accuracy
AI Agent - 信息收集
83.40Thinking Level · Extra High | Tools
79.30Thinking Enabled | Tools
83.20Thinking Enabled | Tools
--
9 additional benchmarks remain in the chart above.

Standard API Pricing: DeepSeek-V4-Pro vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

DeepSeek-V4-Pro
Supplier: DeepSeek-AI
Standard input: $0.435 / 1M tokens
Standard output: $0.87 / 1M tokens
GLM 5.1
Supplier: 智谱AI
Standard input: $1.4 / 1M tokens
Standard output: $4.4 / 1M tokens
Kimi K2.6
Supplier: Facebook AI研究实验室
Standard input: $0.95 / 1M tokens
Standard output: $4 / 1M tokens
Qwen3.8-Max
Supplier: 阿里巴巴
Standard input: ¥12 / 1M tokens
Standard output: ¥36 / 1M tokens
ModelSupplierStandard inputStandard outputBase price applies to
DeepSeek-V4-Pro
DeepSeek-AI$0.435 / 1M tokens$0.87 / 1M tokens
GLM 5.1
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens
Kimi K2.6
Facebook AI研究实验室$0.95 / 1M tokens$4 / 1M tokens
Qwen3.8-Max
阿里巴巴¥12 / 1M tokens¥36 / 1M tokens

Version History

How each version of the DeepSeek-V4-Pro series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkDeepSeek-V4-ProCurrentDeepSeek V3.2DeepSeek-V3.1DeepSeek-R1-0528
HLE
Accuracy
综合评估
48.20Thinking Level · Extra High | Tools
25.10Thinking Enabled
15.90Thinking Enabled
17.70Thinking Enabled
LiveBench
Accuracy
综合评估
71.57Standard Mode
62.20Thinking Enabled
--
--
MMLU Pro
Accuracy
综合评估
87.50Thinking Level · High
--
85.00Thinking Enabled
85.00Thinking Enabled
GPQA Diamond
Accuracy
科学与综合推理
90.10Thinking Level · High
82.40Thinking Enabled
80.10Thinking Enabled
81.00Thinking Enabled
CodeForces
Accuracy
编程与软件工程
3206.00Thinking Level · High
2386.00Thinking Enabled
--
--
LiveCodeBench
Pass @K
编程与软件工程
93.50Thinking Level · High
83.30Thinking Enabled
74.80Thinking Enabled
73.30Thinking Enabled
SWE-Bench Pro - Public
Accuracy
编程与软件工程
55.40Thinking Level · Extra High | Tools
40.90Thinking Enabled
--
--
SWE-bench Verified
Accuracy
编程与软件工程
80.60Thinking Level · Extra High | Tools
73.10Thinking Enabled | Tools
66.00Standard Mode
57.60Thinking Enabled
Creative Writing
Elo、大模型评判两两对战
写作和创作
1552.10Standard Mode
1511.20Standard Mode
1433.20Standard Mode
1418.80Standard Mode
SimpleBench
Score (AVG@5)
常识推理
50.90Standard Mode
--
40.00Standard Mode
40.80Thinking Enabled
BrowseComp
Accuracy
AI Agent - 信息收集
83.40Thinking Level · Extra High | Tools
51.40Thinking Enabled
--
--
Terminal Bench 2.0
Accuracy
AI Agent - 工具使用
67.90Thinking Level · Extra High | Tools
46.40Thinking Enabled | Tools
--
--

Single-Benchmark Version Trend

Viewing: HLE · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the DeepSeek-V4-Pro Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
DeepSeek-V4-Pro
DeepSeek-AI$0.435 / 1M tokens$0.87 / 1M tokens
DeepSeek V3.2
DeepSeek-AI$0.28 / 1M tokens$0.42 / 1M tokens
DeepSeek-V3.1
Fireworks AI$0.56 / 1M tokens$1.68 / 1M tokens
DeepSeek-R1-0528
Fireworks AI$1.35 / 1M tokens$5.4 / 1M tokens

Sources