DataLearner logo

DeepSeek-R1-Distill-Llama-70B Benchmark Details

DeepSeek-R1-Distill-Llama-70B currently shows benchmark results led by MATH-500 (29 / 46, score 94.50), AIME2025 (152 / 216, score 53.70), GPQA Diamond (356 / 463, score 65.20).

Benchmark Results

DeepSeek-R1-Distill-Llama-70B

Benchmark Results

Thinking
Tool usage

General Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
HLE
Thinking Enabled
5.10
480 / 565

Other

2 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
65.20
356 / 463
GPQA Diamond
Thinking Enabled
40.20
433 / 463

Math and Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
MATH-500
Standard Mode
94.50
29 / 46
AIME2025
Thinking Enabled
53.70
152 / 216

Coding and Software Engineer

1 evaluations
Benchmark / mode
Score
Rank/total
LiveCodeBench
Thinking Enabled
26.60
232 / 251

Agent Level Benchmark

2 evaluations
Benchmark / mode
Score
Rank/total
τ²-Bench - Telecom
Thinking EnabledTools
21.90
240 / 264
Terminal Bench Hard
Thinking EnabledTools
1.50
237 / 244

Instruction Following

1 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Thinking Enabled
27.60
272 / 282