DataLearner logo

DeepSeek-R1-0528 Benchmark Details

DeepSeek-R1-0528 currently shows benchmark results led by MMLU-Pro (28 / 175, score 85), MATH-500 (8 / 46, score 98), AIME 2024 (13 / 61, score 91.40).

Benchmark Results

DeepSeek-R1-0528

Benchmark Results

Thinking
Tool usage

Knowledge Exams

2 evaluations
Benchmark / mode
Score
Rank/total
MMLU-Pro
Thinking Mode
85
28 / 175
HLE
Thinking Mode
17.70
169 / 235

Abstract Generalization

2 evaluations
Benchmark / mode
Score
Rank/total
ARC-AGI-1
Thinking Mode
21.20
157 / 176
ARC-AGI-2
Thinking Mode
1.30
151 / 164

Scientific Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Mode
81
122 / 253
CritPt
Thinking Mode
1.40
139 / 204

Honesty & Factuality

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleQA
Thinking Mode
27.80
27 / 46

Repository Engineering

1 evaluations
Benchmark / mode
Score
Rank/total
SWE-bench Verified
Thinking Mode
57.60
86 / 116

Mathematics

5 evaluations
Benchmark / mode
Score
Rank/total
MATH-500
Thinking Mode
98
8 / 46
AIME 2024
Thinking Mode
91.40
13 / 61
AIME2025
Thinking Mode
87.50
49 / 110
IMO-ProofBench
Thinking Mode
29
7 / 16
3.80
22 / 24

Algorithmic Coding

1 evaluations
Benchmark / mode
Score
Rank/total
LiveCodeBench
Thinking Mode
73.30
49 / 125

Writing

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1421.10
63 / 110

Agentic Development

3 evaluations
Benchmark / mode
Score
Rank/total
Aider-Polyglot
Thinking Mode
71.40
15 / 59
Terminal Bench Hard
Thinking ModeTools
15.90
165 / 244
Terminal-Bench
Thinking Mode
5.70
35 / 35

Commonsense

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Thinking Mode
40.80
68 / 96

Capability Indices

1 evaluations
Benchmark / mode
Score
Rank/total
ECI
unknown
141.30
91 / 167