DeepSeek-V4-ProvsDeepSeek V3.2
Across 10 shared benchmarks, DeepSeek-V4-Pro leads overall: DeepSeek-V4-Pro wins 5, DeepSeek V3.2 wins 4, with 1 ties and an average score difference of +7.31.
DeepSeek-V4-Pro
DeepSeek-AI · 2026-08-13 · Reasoning model
DeepSeek V3.2
DeepSeek-AI · 2025-12-01 · Reasoning model
DeepSeek-V4-Pro5 wins(50%)Ties1(40%)4 winsDeepSeek V3.2
Benchmark scores
Grouped by capability, sorted by largest gap within each. 10 shared benchmarks.
General Knowledge
Even 3/3| Benchmark | DeepSeek-V4-Pro | DeepSeek V3.2 | Diff |
|---|---|---|---|
| LiveBench | 71.5732 / 117Normal (No Tools) | 51.8489 / 117Normal (No Tools) | +19.73 |
| HLE | 8.20416 / 563Normal (No Tools) · Text only | 11.20375 / 563Normal (No Tools) · Text only | -3 |
| CritPt | 0.90158 / 200Normal (No Tools) | 0.90158 / 200Normal (No Tools) | — |
Agent Level Benchmark
DeepSeek-V4-Pro 2/2| Benchmark | DeepSeek-V4-Pro | DeepSeek V3.2 | Diff |
|---|---|---|---|
| τ²-Bench - Telecom | 91.2059 / 264Normal (With Tools) | 78.90117 / 264Normal (With Tools) | +12.30 |
| Terminal Bench Hard | 36.4067 / 244Normal (With Tools) | 32.6093 / 244Normal (With Tools) | +3.80 |
Coding and Software Engineer
DeepSeek V3.2 1/1| Benchmark | DeepSeek-V4-Pro | DeepSeek V3.2 | Diff |
|---|---|---|---|
| LiveCodeBench | 56.80142 / 250Normal (No Tools) | 59.30133 / 250Normal (No Tools) | -2.50 |
General Evaluation
DeepSeek V3.2 1/1| Benchmark | DeepSeek-V4-Pro | DeepSeek V3.2 | Diff |
|---|---|---|---|
| GPQA Diamond | 72.90301 / 462Normal (No Tools) | 75.10279 / 462Normal (No Tools) | -2.20 |
Instruction Following
DeepSeek V3.2 1/1| Benchmark | DeepSeek-V4-Pro | DeepSeek V3.2 | Diff |
|---|---|---|---|
| IF Bench | 45.80176 / 282Normal (No Tools) | 49161 / 282Normal (No Tools) | -3.20 |
Long Context
DeepSeek-V4-Pro 1/1| Benchmark | DeepSeek-V4-Pro | DeepSeek V3.2 | Diff |
|---|---|---|---|
| AA-LCR | 53135 / 170Normal (No Tools) | 45.70144 / 170Normal (No Tools) | +7.30 |
Writing and Creative Capabilities
DeepSeek-V4-Pro 1/1| Benchmark | DeepSeek-V4-Pro | DeepSeek V3.2 | Diff |
|---|---|---|---|
| Creative Writing | 1,55247 / 106Normal (No Tools) | 1,51149 / 106Normal (No Tools) | +40.90 |
Specs
| Field | DeepSeek-V4-Pro | DeepSeek V3.2 |
|---|---|---|
| Publisher | DeepSeek-AI | DeepSeek-AI |
| Release date | 2026-08-13 | 2025-12-01 |
| Model type | Reasoning model | Reasoning model |
| Architecture | MoE | MoE |
| Parameters | 1.6T | 671B |
| Context length | 1M | 128K |
| Max output | 384K | 8K |
API pricing
Prices use DataLearner records when available; missing fields are not inferred.
| Item | DeepSeek-V4-Pro | DeepSeek V3.2 |
|---|---|---|
| Text input | $0.435 / 1M tokens | $0.28 / 1M tokens |
| Text output | $0.87 / 1M tokens | $0.42 / 1M tokens |
| Cache read | $0.003625 / 1M tokens | $0.028 / 1M tokens |
| Cache write | Not public | $0.28 / 1M tokens |
Summary
- DeepSeek-V4-Proleads in:Agent Level Benchmark (2/2), Long Context (1/1), Writing and Creative Capabilities (1/1)
- DeepSeek V3.2leads in:Coding and Software Engineer (1/1), General Evaluation (1/1), Instruction Following (1/1)
- Tied in:General Knowledge
On average across the 10 shared benchmarks, DeepSeek-V4-Pro scores 7.31 higher.
Largest single-benchmark gap: Creative Writing — DeepSeek-V4-Pro 1,552 vs DeepSeek V3.2 1,511 (+40.90).
Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.