DataLearner logo

DeepSeek V3.2 Speciale Benchmark Details

DeepSeek V3.2 Speciale currently shows benchmark results led by LiveCodeBench (11 / 250, score 89.60), AIME2025 (25 / 215, score 96), GPQA Diamond (132 / 462, score 87.10).

Benchmark Results

DeepSeek V3.2 Speciale

Benchmark Results

Thinking
Tool usage

General Knowledge

3 evaluations
Benchmark / mode
Score
Rank/total
HLE
Thinking Mode
30.60
197 / 563
HLE
Thinking Mode
28.70
216 / 563
CritPt
Thinking Mode
7.40
85 / 200

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Mode
87.10
132 / 462

Coding and Software Engineer

2 evaluations
Benchmark / mode
Score
Rank/total
CodeForces
Thinking Mode
2701
8 / 21
LiveCodeBench
Thinking Mode
89.60
11 / 250

Math and Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
AIME2025
Thinking Mode
96
25 / 215

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1273.40
75 / 106

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Standard Mode
52.60
48 / 93

Instruction Following

1 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Thinking Mode
63.90
108 / 282

Agent Level Benchmark

1 evaluations
Benchmark / mode
Score
Rank/total
Terminal Bench Hard
Thinking ModeTools
34.80
75 / 244