DataLearner logo

DeepSeek-V3.1 Benchmark Details

DeepSeek-V3.1 currently shows benchmark results led by MMLU (1 / 66, score 93.40), SimpleQA (4 / 47, score 93.40), AIME 2024 (7 / 62, score 93.10).

Benchmark Results

DeepSeek-V3.1

Benchmark Results

Thinking

General Knowledge

7 evaluations
Benchmark / mode
Score
Rank/total
93.40
1 / 66
91.80
4 / 66
85
26 / 132
83.70
43 / 132
80.10
81 / 187
74.90
101 / 187
15.90
133 / 172

Common Sense

1 evaluations
Benchmark / mode
Score
Rank/total
93.40
4 / 47

Coding and Software Engineer

3 evaluations
Benchmark / mode
Score
Rank/total
74.80
41 / 123
56.40
79 / 123

Math and Reasoning

4 evaluations
Benchmark / mode
Score
Rank/total
93.10
7 / 62
66.30
40 / 62
88.40
43 / 107
49.80
88 / 107

AI Agent - Tool Usage

1 evaluations
Benchmark / mode
Score
Rank/total
31.30
19 / 35

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
Simple Bench
Standard Mode
40
40 / 63