DataLearner logo

DeepSeek V3.2vsDeepSeek-V3.1

Across 6 shared benchmarks, DeepSeek V3.2 leads overall: DeepSeek V3.2 wins 5, DeepSeek-V3.1 wins 1, with 0 ties and an average score difference of +15.60.

DeepSeek-AI
DeepSeek V3.2

DeepSeek-AI · 2025-12-01 · Reasoning model

DeepSeek-AI
DeepSeek-V3.1

DeepSeek-AI · 2025-08-20 · Chat model

DeepSeek V3.25 wins(83%)(17%)1 winDeepSeek-V3.1

Benchmark scores

Grouped by capability, sorted by largest gap within each. 6 shared benchmarks.

Coding and Software Engineer

DeepSeek V3.2 1/1
BenchmarkDeepSeek V3.2DeepSeek-V3.1Diff
LiveCodeBench59.30133 / 250Normal (No Tools)56.40145 / 250Normal (No Tools)+2.90

General Evaluation

DeepSeek V3.2 1/1
BenchmarkDeepSeek V3.2DeepSeek-V3.1Diff
GPQA Diamond75.10279 / 462Normal (No Tools)74.90283 / 462Normal (No Tools)+0.20

General Knowledge

DeepSeek V3.2 1/1
BenchmarkDeepSeek V3.2DeepSeek-V3.1Diff
HLE11.20375 / 563Normal (No Tools) · Text only6.70449 / 563Normal (No Tools) · Text only+4.50

Long Context

DeepSeek-V3.1 1/1
BenchmarkDeepSeek V3.2DeepSeek-V3.1Diff
AA-LCR45.70144 / 170Normal (No Tools)47140 / 170Normal (No Tools)-1.30

Math and Reasoning

DeepSeek V3.2 1/1
BenchmarkDeepSeek V3.2DeepSeek-V3.1Diff
AIME202559139 / 215Normal (No Tools)49.70154 / 215Normal (No Tools)+9.30

Writing and Creative Capabilities

DeepSeek V3.2 1/1
BenchmarkDeepSeek V3.2DeepSeek-V3.1Diff
Creative Writing1,51149 / 106Normal (No Tools)1,43357 / 106Normal (No Tools)+78

Specs

FieldDeepSeek V3.2DeepSeek-V3.1
PublisherDeepSeek-AIDeepSeek-AI
Release date2025-12-012025-08-20
Model typeReasoning modelChat model
ArchitectureMoEMoE
Parameters671B671B
Context length128K128K
Max output8K8K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemDeepSeek V3.2DeepSeek-V3.1
Text input$0.28 / 1M tokens$0.56 / 1M tokens
Text output$0.42 / 1M tokens$1.68 / 1M tokens
Cache read$0.028 / 1M tokens$0.28 / 1M tokens
Cache write$0.28 / 1M tokens$0.56 / 1M tokens

Summary

  • DeepSeek V3.2leads in:Coding and Software Engineer (1/1), General Evaluation (1/1), General Knowledge (1/1), Math and Reasoning (1/1), Writing and Creative Capabilities (1/1)
  • DeepSeek-V3.1leads in:Long Context (1/1)

On average across the 6 shared benchmarks, DeepSeek V3.2 scores 15.60 higher.

Largest single-benchmark gap: Creative Writing — DeepSeek V3.2 1,511 vs DeepSeek-V3.1 1,433 (+78).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.