DataLearner logo

DeepSeek-V3.1 Benchmark Details

DeepSeek-V3.1 currently shows benchmark results led by MMLU (1 / 124, score 93.40), SimpleQA (4 / 47, score 93.40), AIME 2024 (7 / 62, score 93.10).

Benchmark Results

DeepSeek-V3.1

Benchmark Results

Thinking
Tool usage

General Knowledge

5 evaluations
Benchmark / mode
Score
Rank/total
MMLU
Standard Mode
91.80
4 / 124
MMLU
Thinking Mode
93.40
1 / 124
MMLU Pro
Standard Mode
83.70
44 / 134
MMLU Pro
Thinking Mode
85
27 / 134
HLE
Thinking Mode
15.90
156 / 197

Other

2 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
74.90
175 / 274
GPQA Diamond
Thinking Mode
80.10
146 / 274

Common Sense

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleQA
Thinking Mode
93.40
4 / 47

Coding and Software Engineer

3 evaluations
Benchmark / mode
Score
Rank/total
LiveCodeBench
Standard Mode
56.40
84 / 128
LiveCodeBench
Thinking Mode
74.80
44 / 128
SWE-bench Verified
Standard Mode
66
76 / 116

Math and Reasoning

4 evaluations
Benchmark / mode
Score
Rank/total
AIME 2024
Standard Mode
66.30
40 / 62
AIME 2024
Thinking Mode
93.10
7 / 62
AIME2025
Standard Mode
49.80
88 / 107
AIME2025
Thinking Mode
88.40
42 / 107

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1433.20
57 / 106

AI Agent - Tool Usage

1 evaluations
Benchmark / mode
Score
Rank/total
Terminal-Bench
Standard ModeTools
31.30
19 / 35

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Standard Mode
40
68 / 94