DataLearner logo
DE

DeepSeek-R1-Distill-Llama-70B

Reasoning modelDeepSeek R1 DistillDeepSeek R1 Distill

DeepSeek-R1-Distill-Llama-70B

Release date: 2025-01-20Updated: 2025-02-08Views: 1,555
Parameters
70B
Context length
128K
Multilingual
No data
Reasoning ability
No data

DeepSeek-R1-Distill-Llama-70B is an AI model published by DeepSeek-AI, released on 2025-01-20, for Reasoning model, with 70B parameters, and 128K context length, requiring about 140GB storage, with a 94.50 score on MATH-500.

Data sourced primarily from official releases (GitHub, Hugging Face, papers), then benchmark leaderboards, then third-party evaluators. Learn about our data methodology

DeepSeek-R1-Distill-Llama-70B

Model basics

Reasoning traces
Supported
Thinking modes
Thinking modes not supported
Context length
128K tokens
Max output length
No data
Model type
Reasoning model
Modality (in / out)
No data
Release date
2025-01-20
Model file size
140GB
MoE architecture
No
Total params / Active params
70B / Not applicable
Knowledge cutoff
No data
DeepSeek-R1-Distill-Llama-70B

Open source & experience

Code license
Weights license
MIT License- Commercial use permitted
Live demo
N/A
DeepSeek-R1-Distill-Llama-70B

Official resources

Paper
N/A
DataLearnerAI blog
N/A
DeepSeek-R1-Distill-Llama-70B

API details

API speed
No data
💡Default unit: $/1M tokens. If vendors use other units, follow their published pricing.
Standard
TypeConditionInputOutput
Text-$0.600/ 1M tokens$1.20/ 1M tokens

“—” means the modality is not billed in that direction, or the vendor has not published a price for it.

DeepSeek-R1-Distill-Llama-70B

Benchmark Results

DeepSeek-R1-Distill-Llama-70B currently shows benchmark results led by MATH-500 (28 / 45, score 94.50), AIME2025 (151 / 215, score 53.70), GPQA Diamond (355 / 462, score 65.20). This page also consolidates core specs, context limits, and API pricing so you can evaluate the model from benchmark results and deployment constraints together.

Thinking
Tool usage

General Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
HLE
Thinking Mode
5.10
478 / 563

General Evaluation

2 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
65.20
355 / 462
GPQA Diamond
Thinking Mode
40.20
432 / 462

Math and Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
MATH-500
Standard Mode
94.50
28 / 45
AIME2025
Thinking Mode
53.70
151 / 215

Coding and Software Engineer

1 evaluations
Benchmark / mode
Score
Rank/total
LiveCodeBench
Thinking Mode
26.60
231 / 250

Agent Level Benchmark

2 evaluations
Benchmark / mode
Score
Rank/total
τ²-Bench - Telecom
Thinking ModeTools
21.90
240 / 264
Terminal Bench Hard
Thinking ModeTools
1.50
237 / 244

Instruction Following

1 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Thinking Mode
27.60
272 / 282
DeepSeek-R1-Distill-Llama-70B

Publisher

DeepSeek-R1-Distill-Llama-70B

Model Overview

DeepSeek-R1-Distill-Llama-70B is an AI model published by DeepSeek-AI, released on 2025-01-20, for Reasoning model, with 70B parameters, and 128K context length, requiring about 140GB storage, with a 94.50 score on MATH-500.

DataLearner on WeChat

Follow DataLearner on WeChat for AI model updates and research notes.

DataLearner WeChat QR code