DataLearner logo

DeepSeek-R1 Benchmark Details

DeepSeek-R1 currently shows benchmark results led by MMLU (8 / 124, score 90.80), MMLU Pro (40 / 176, score 84), MATH-500 (13 / 45, score 97.30).

Benchmark Results

DeepSeek-R1

Benchmark Results

Thinking
Tool usage

General Knowledge

5 evaluations
Benchmark / mode
Score
Rank/total
MMLU
Standard Mode
90.80
8 / 124
MMLU Pro
Standard Mode
84
40 / 176
ARC-AGI-1
Standard Mode
15.80
135 / 147
HLE
Thinking Enabled
8.50
412 / 563
CritPt
Thinking Enabled
0.60
168 / 200

Other

2 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
71.50
306 / 462
GPQA Diamond
Thinking Enabled
70.80
317 / 462

Common Sense

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleQA
Standard Mode
30.10
24 / 47

Coding and Software Engineer

5 evaluations
Benchmark / mode
Score
Rank/total
LiveCodeBench
Standard Mode
65.90
107 / 250
LiveCodeBench
Thinking Enabled
61.70
126 / 250
SWE-bench Verified
Standard Mode
49.20
98 / 116
SciCode
Thinking Enabled
38.30
110 / 130
WeirdML v2
Standard ModeTools
36.49
49 / 52

Math and Reasoning

5 evaluations
Benchmark / mode
Score
Rank/total
MATH-500
Standard Mode
97.30
13 / 45
AIME 2024
Standard Mode
79.80
28 / 62
AIME2025
Standard Mode
70
120 / 215
AIME2025
Thinking Enabled
68
124 / 215
IMO-ProofBench Advanced
Thinking Enabled
3.80
22 / 24

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1500
51 / 106

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Standard Mode
30.90
75 / 93

Agent Level Benchmark

6 evaluations
Benchmark / mode
Score
Rank/total
Aider-Polyglot
Thinking Enabled
56.90
25 / 59
BALROG
Standard ModeTools
34.90
8 / 12
METR Time Horizons v1.1
Standard ModeTools
26.93
21 / 22
τ²-Bench - Telecom
Thinking EnabledTools
11.40
262 / 264
τ³-Banking
Thinking EnabledTools
6.40
144 / 164
Terminal Bench Hard
Thinking EnabledTools
6.10
204 / 244

Instruction Following

1 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Thinking Enabled
39
223 / 282

AI Agent - Tool Usage

1 evaluations
Benchmark / mode
Score
Rank/total
Terminal-Bench 2.1
Thinking EnabledTools
19.10
163 / 192

Other

1 evaluations
Benchmark / mode
Score
Rank/total
Fiction.liveBench
Standard Mode
69.40
11 / 16

Multimodal Understanding

1 evaluations
Benchmark / mode
Score
Rank/total
GDP.pdf
Thinking Enabled
2.40
105 / 118