DataLearner logo

DeepSeek V3.2 Benchmark Details

DeepSeek V3.2 currently shows benchmark results led by LiveCodeBench (24 / 128, score 83.30), AIME2025 (29 / 107, score 93.10), τ²-Bench (14 / 44, score 80.30). This page also tracks comparisons against 3 predecessor or same-series models. 1 source link is attached for reference.

Benchmark Results

DeepSeek V3.2

Benchmark Results

Thinking
Tool usage

General Knowledge

5 evaluations
Benchmark / mode
Score
Rank/total
LiveBench
Standard Mode
51.84
89 / 117
LiveBench
Thinking Mode
62.20
60 / 117
ARC-AGI-1
Thinking Mode
57
65 / 92
HLE
Thinking Mode
25.10
124 / 197
ARC-AGI-2
Thinking Mode
4
74 / 85

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Mode
82.40
128 / 274

Coding and Software Engineer

5 evaluations
Benchmark / mode
Score
Rank/total
CodeForces
Thinking Mode
2386
12 / 21
LiveCodeBench
Thinking Mode
83.30
24 / 128
SWE-bench Verified
Thinking Mode
70.20
62 / 116
SWE-bench Verified
Thinking ModeTools
73.10
50 / 116
40.90
56 / 62

Math and Reasoning

3 evaluations
Benchmark / mode
Score
Rank/total
AIME2025
Thinking Mode
93.10
29 / 107
AIME 2026
Thinking Mode
92.70
10 / 21
2.10
56 / 80

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1511.20
49 / 106

Agent Level Benchmark

1 evaluations
Benchmark / mode
Score
Rank/total
τ²-Bench
Thinking ModeTools
80.30
14 / 44

AI Agent - Information Search

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
Thinking Mode
51.40
44 / 57

AI Agent - Tool Usage

1 evaluations
Benchmark / mode
Score
Rank/total
Terminal Bench 2.0
Thinking ModeTools
46.40
41 / 48

Claw-style Agent Evaluation

2 evaluations
Benchmark / mode
Score
Rank/total
Pinch Bench
Thinking ModeTools
84.30
19 / 38
Claw Bench
Thinking ModeTools
79
21 / 29

Version History

How each version of the DeepSeek V3.2 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

8 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkDeepSeek V3.2CurrentDeepSeek-V3.1DeepSeek-V3-0324DeepSeek-V3
ARC-AGI-1
Score (%)
综合评估
57.00Thinking Enabled
--
9.00Standard Mode
--
HLE
Accuracy
综合评估
25.10Thinking Enabled
15.90Thinking Enabled
5.20Standard Mode
--
GPQA Diamond
Accuracy
科学与综合推理
82.40Thinking Enabled
80.10Thinking Enabled
68.40Standard Mode
59.10Standard Mode
LiveCodeBench
Pass @K
编程与软件工程
83.30Thinking Enabled
74.80Thinking Enabled
49.20Standard Mode
34.60Standard Mode
SWE-bench Verified
Accuracy
编程与软件工程
73.10Thinking Enabled | Tools
66.00Standard Mode
38.80Standard Mode
--
AIME2025
Accuracy
数学推理
93.10Thinking Enabled
88.40Thinking Enabled
47.70Standard Mode
--
Creative Writing
Elo、大模型评判两两对战
写作和创作
1511.20Standard Mode
1433.20Standard Mode
1470.20Standard Mode
--
τ²-Bench
Accuracy
Agent能力评测
80.30Thinking Enabled | Tools
--
38.80Standard Mode | Tools
--

Single-Benchmark Version Trend

Viewing: ARC-AGI-1 · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the DeepSeek V3.2 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
DeepSeek V3.2
DeepSeek-AI$0.28 / 1M tokens$0.42 / 1M tokens
DeepSeek-V3.1
Fireworks AI$0.56 / 1M tokens$1.68 / 1M tokens
DeepSeek-V3-0324
DeepInfra$0.2 / 1M tokens$0.88 / 1M tokens
DeepSeek-V3
DeepSeek-AI$0.27 / 1M tokens$1.1 / 1M tokens

Sources