DataLearner logo

DeepSeek-V4-ProvsDeepSeek V3.2

Across 10 shared benchmarks, DeepSeek-V4-Pro leads overall: DeepSeek-V4-Pro wins 5, DeepSeek V3.2 wins 4, with 1 ties and an average score difference of +7.31.

DeepSeek-AI
DeepSeek-V4-Pro

DeepSeek-AI · 2026-08-13 · Reasoning model

DeepSeek-AI
DeepSeek V3.2

DeepSeek-AI · 2025-12-01 · Reasoning model

DeepSeek-V4-Pro5 wins(50%)Ties1(40%)4 winsDeepSeek V3.2

Benchmark scores

Grouped by capability, sorted by largest gap within each. 10 shared benchmarks.

General Knowledge

Even 3/3
BenchmarkDeepSeek-V4-ProDeepSeek V3.2Diff
LiveBench71.5732 / 117Normal (No Tools)51.8489 / 117Normal (No Tools)+19.73
HLE8.20416 / 563Normal (No Tools) · Text only11.20375 / 563Normal (No Tools) · Text only-3
CritPt0.90158 / 200Normal (No Tools)0.90158 / 200Normal (No Tools)

Agent Level Benchmark

DeepSeek-V4-Pro 2/2
BenchmarkDeepSeek-V4-ProDeepSeek V3.2Diff
τ²-Bench - Telecom91.2059 / 264Normal (With Tools)78.90117 / 264Normal (With Tools)+12.30
Terminal Bench Hard36.4067 / 244Normal (With Tools)32.6093 / 244Normal (With Tools)+3.80

Coding and Software Engineer

DeepSeek V3.2 1/1
BenchmarkDeepSeek-V4-ProDeepSeek V3.2Diff
LiveCodeBench56.80142 / 250Normal (No Tools)59.30133 / 250Normal (No Tools)-2.50

General Evaluation

DeepSeek V3.2 1/1
BenchmarkDeepSeek-V4-ProDeepSeek V3.2Diff
GPQA Diamond72.90301 / 462Normal (No Tools)75.10279 / 462Normal (No Tools)-2.20

Instruction Following

DeepSeek V3.2 1/1
BenchmarkDeepSeek-V4-ProDeepSeek V3.2Diff
IF Bench45.80176 / 282Normal (No Tools)49161 / 282Normal (No Tools)-3.20

Long Context

DeepSeek-V4-Pro 1/1
BenchmarkDeepSeek-V4-ProDeepSeek V3.2Diff
AA-LCR53135 / 170Normal (No Tools)45.70144 / 170Normal (No Tools)+7.30

Writing and Creative Capabilities

DeepSeek-V4-Pro 1/1
BenchmarkDeepSeek-V4-ProDeepSeek V3.2Diff
Creative Writing1,55247 / 106Normal (No Tools)1,51149 / 106Normal (No Tools)+40.90

Specs

FieldDeepSeek-V4-ProDeepSeek V3.2
PublisherDeepSeek-AIDeepSeek-AI
Release date2026-08-132025-12-01
Model typeReasoning modelReasoning model
ArchitectureMoEMoE
Parameters1.6T671B
Context length1M128K
Max output384K8K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemDeepSeek-V4-ProDeepSeek V3.2
Text input$0.435 / 1M tokens$0.28 / 1M tokens
Text output$0.87 / 1M tokens$0.42 / 1M tokens
Cache read$0.003625 / 1M tokens$0.028 / 1M tokens
Cache writeNot public$0.28 / 1M tokens

Summary

  • DeepSeek-V4-Proleads in:Agent Level Benchmark (2/2), Long Context (1/1), Writing and Creative Capabilities (1/1)
  • DeepSeek V3.2leads in:Coding and Software Engineer (1/1), General Evaluation (1/1), Instruction Following (1/1)
  • Tied in:General Knowledge

On average across the 10 shared benchmarks, DeepSeek-V4-Pro scores 7.31 higher.

Largest single-benchmark gap: Creative Writing — DeepSeek-V4-Pro 1,552 vs DeepSeek V3.2 1,511 (+40.90).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.