DataLearner logo

DeepSeek-V4-Pro Benchmark Details

DeepSeek-V4-Pro currently shows benchmark results led by LiveCodeBench (1 / 123, score 93.50), MMLU Pro (11 / 132, score 87.50), SWE-bench Verified (11 / 112, score 80.60). This page also compares it with 2 competitor models and 3 predecessor or same-series models, including performance and pricing views when available. 1 source link is attached for reference.

Benchmark Results

DeepSeek-V4-Pro

Benchmark Results

Thinking
Tool usage

General Knowledge

12 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
72.90
108 / 187
89.10
27 / 187
90.10
23 / 187
MMLU Pro
Standard Mode
82.90
48 / 132
87.10
13 / 132
87.50
11 / 132
LiveBench
Standard Mode
73.58
23 / 115
HLE
Standard Mode
7.70
156 / 172
HLE
High
34.50
77 / 172
HLE
HighTools
44.70
44 / 172
HLE
Max
37.70
67 / 172
HLE
Thinking Level · Extra HighTools
48.20
36 / 172

Coding and Software Engineer

14 evaluations
Benchmark / mode
Score
Rank/total
2919
4 / 16
3206
2 / 16
LiveCodeBench
Standard Mode
56.80
76 / 123
89.80
6 / 123
93.50
1 / 123
SWE-bench Verified
Standard ModeTools
73.60
45 / 112
79.40
19 / 112
SWE-bench Verified
Thinking Level · Extra HighTools
80.60
11 / 112
SWE-bench Multilingual
Standard ModeTools
69.80
18 / 23
74.10
8 / 23
SWE-bench Multilingual
Thinking Level · Extra HighTools
76.20
6 / 23
SWE-Bench Pro - Public
Standard ModeTools
52.10
38 / 54
54.40
29 / 54
SWE-Bench Pro - Public
Thinking Level · Extra HighTools
55.40
26 / 54

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
Simple Bench
Standard Mode
50.90
28 / 63

AI Agent - Information Search

2 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
HighTools
80.40
16 / 53
BrowseComp
Thinking Level · Extra HighTools
83.40
13 / 53

AI Agent - Tool Usage

3 evaluations
Benchmark / mode
Score
Rank/total
Terminal Bench 2.0
Standard ModeTools
59.10
22 / 47
63.30
14 / 47
Terminal Bench 2.0
Thinking Level · Extra HighTools
67.90
9 / 47

Math and Reasoning

3 evaluations
Benchmark / mode
Score
Rank/total
IMO-AnswerBench
Standard Mode
35.30
21 / 21
88
6 / 21
89.80
4 / 21

Productivity Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA
Thinking Level · Extra HighTools
1554
4 / 21

Competitor Comparison

Benchmark scores for DeepSeek-V4-Pro compared against top models in its class

DeepSeek-V4-ProGLM 5.1Kimi K2.6
Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

10 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkDeepSeek-V4-ProCurrentGLM 5.1Kimi K2.6
GPQA Diamond
综合评估
90.10Thinking Level · High
86.20Thinking Enabled
90.50Thinking Enabled
HLE
综合评估
48.20Thinking Level · Extra High | Tools
52.30Thinking Enabled | Tools
54.00Thinking Enabled | Tools
LiveBench
综合评估
73.58Standard Mode
70.18Standard Mode
72.17Thinking Enabled
LiveCodeBench
编程与软件工程
93.50Thinking Level · High
--
89.60Thinking Enabled
SWE-bench Multilingual
编程与软件工程
76.20Thinking Level · Extra High | Tools
--
76.70Thinking Enabled | Tools
SWE-Bench Pro - Public
编程与软件工程
55.40Thinking Level · Extra High | Tools
58.40Thinking Enabled | Tools
58.60Thinking Enabled | Tools
SWE-bench Verified
编程与软件工程
80.60Thinking Level · Extra High | Tools
--
80.20Thinking Enabled | Tools
BrowseComp
AI Agent - 信息收集
83.40Thinking Level · Extra High | Tools
79.30Thinking Enabled | Tools
83.20Thinking Enabled | Tools
Terminal Bench 2.0
AI Agent - 工具使用
67.90Thinking Level · Extra High | Tools
63.50Thinking Enabled | Tools
66.70Thinking Enabled | Tools
IMO-AnswerBench
数学推理
89.80Thinking Level · High
83.80Thinking Enabled
86.00Thinking Enabled

Standard API Pricing: DeepSeek-V4-Pro vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
DeepSeek-V4-Pro
DeepSeek-AI$0.435 / 1M tokens$0.87 / 1M tokens
GLM 5.1
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens
Kimi K2.6
Facebook AI研究实验室$0.95 / 1M tokens$4 / 1M tokens

Version History

How each version of the DeepSeek-V4-Pro series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

11 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkDeepSeek-V4-ProCurrentDeepSeek V3.2DeepSeek-V3.1DeepSeek-R1-0528
GPQA Diamond
综合评估
90.10Thinking Level · High
82.40Thinking Enabled
80.10Thinking Enabled
81.00Thinking Enabled
HLE
综合评估
48.20Thinking Level · Extra High | Tools
25.10Thinking Enabled
15.90Thinking Enabled
17.70Thinking Enabled
LiveBench
综合评估
73.58Standard Mode
62.20Thinking Enabled
--
--
MMLU Pro
综合评估
87.50Thinking Level · High
--
85.00Thinking Enabled
85.00Thinking Enabled
CodeForces
编程与软件工程
3206.00Thinking Level · High
2386.00Thinking Enabled
--
--
LiveCodeBench
编程与软件工程
93.50Thinking Level · High
83.30Thinking Enabled
74.80Thinking Enabled
73.30Thinking Enabled
SWE-Bench Pro - Public
编程与软件工程
55.40Thinking Level · Extra High | Tools
40.90Thinking Enabled
--
--
SWE-bench Verified
编程与软件工程
80.60Thinking Level · Extra High | Tools
73.10Thinking Enabled | Tools
66.00Standard Mode
57.60Thinking Enabled
Simple Bench
常识推理
50.90Standard Mode
--
40.00Standard Mode
40.80Thinking Enabled
BrowseComp
AI Agent - 信息收集
83.40Thinking Level · Extra High | Tools
51.40Thinking Enabled
--
--
Terminal Bench 2.0
AI Agent - 工具使用
67.90Thinking Level · Extra High | Tools
46.40Thinking Enabled | Tools
--
--

Single-Benchmark Version Trend

Viewing: GPQA Diamond · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the DeepSeek-V4-Pro Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
DeepSeek-V4-Pro
DeepSeek-AI$0.435 / 1M tokens$0.87 / 1M tokens
DeepSeek V3.2
DeepSeek-AI$0.28 / 1M tokens$0.42 / 1M tokens
DeepSeek-V3.1
Fireworks AI$0.56 / 1M tokens$1.68 / 1M tokens
DeepSeek-R1-0528
Fireworks AI$1.35 / 1M tokens$5.4 / 1M tokens

Sources