DataLearner logo

DeepSeek V3.2-Exp Benchmark Details

DeepSeek V3.2-Exp currently shows benchmark results led by SimpleQA (1 / 47, score 97.10), Aider-Polyglot (11 / 59, score 74.20), MMLU Pro (26 / 132, score 85).

Benchmark Results

DeepSeek V3.2-Exp

Benchmark Results

Thinking

General Knowledge

9 evaluations
Benchmark / mode
Score
Rank/total
85
26 / 132
84
38 / 132
79.90
84 / 188
74
103 / 188
LiveBench
Standard Mode
49.85
91 / 115
LiveBench
Thinking Mode
58.90
73 / 115
20.30
118 / 173
19.80
120 / 173
8.60
153 / 173

Common Sense

1 evaluations
Benchmark / mode
Score
Rank/total
97.10
1 / 47

Coding and Software Engineer

3 evaluations
Benchmark / mode
Score
Rank/total
74.10
42 / 123
55
85 / 123
67.80
72 / 113

Math and Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
89.30
39 / 107
58
84 / 107

AI Agent - Tool Usage

2 evaluations
Benchmark / mode
Score
Rank/total
37.70
14 / 35

Agent Level Benchmark

5 evaluations
Benchmark / mode
Score
Rank/total
Aider-Polyglot
Standard Mode
70.20
17 / 59
Aider-Polyglot
Thinking Mode
74.20
11 / 59
66.70
27 / 43

Instruction Following

1 evaluations
Benchmark / mode
Score
Rank/total
54.10
28 / 31

AI Agent - Information Search

1 evaluations
Benchmark / mode
Score
Rank/total
40.10
48 / 53