DataLearner logo
DE

DeepSeek-R1-Distill-Qwen-7B

Reasoning modelDeepSeek R1 DistillDeepSeek R1 Distill

DeepSeek-R1-Distill-Qwen-7B

Release date: 2025-01-20Updated: 2025-02-27Views: 1,359
Live demoGitHubHugging FaceCompare
Parameters
7B
Context length
128K
Multilingual
Supported
Reasoning ability
No data

DeepSeek-R1-Distill-Qwen-7B is an AI model published by DeepSeek-AI, released on 2025-01-20, for Reasoning model, with 7B parameters, and 128K context length, requiring about 14GB storage, with a 91.40 score on MATH-500.

Data sourced primarily from official releases (GitHub, Hugging Face, papers), then benchmark leaderboards, then third-party evaluators. Learn about our data methodology

DeepSeek-R1-Distill-Qwen-7B

Model basics

Reasoning traces
Supported
Thinking modes
Thinking modes not supported
Context length
128K tokens
Max output length
No data
Model type
Reasoning model
Modality (in / out)
No data
Release date
2025-01-20
Model file size
14GB
MoE architecture
No
Total params / Active params
7B / Not applicable
Knowledge cutoff
No data
DeepSeek-R1-Distill-Qwen-7B

Open source & experience

Code license
Weights license
MIT License- Commercial use permitted
GitHub repo
N/A
Live demo
N/A
DeepSeek-R1-Distill-Qwen-7B

Official resources

Paper
DataLearnerAI blog
N/A
DeepSeek-R1-Distill-Qwen-7B

API details

API speed
No data
No public API pricing yet.
DeepSeek-R1-Distill-Qwen-7B

Benchmark Results

DeepSeek-R1-Distill-Qwen-7B currently shows benchmark results led by AIME 2024 (45 / 62, score 53.30), MATH-500 (33 / 45, score 91.40), GPQA Diamond (247 / 274, score 49.50). This page also consolidates core specs, context limits, and API pricing so you can evaluate the model from benchmark results and deployment constraints together.

Thinking

General Evaluation

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
49.50
247 / 274

Math and Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
MATH-500
Standard Mode
91.40
33 / 45
AIME 2024
Standard Mode
53.30
45 / 62
DeepSeek-R1-Distill-Qwen-7B

Publisher

DeepSeek-R1-Distill-Qwen-7B

Model Overview

DeepSeek-R1-Distill-Qwen-7B is an AI model published by DeepSeek-AI, released on 2025-01-20, for Reasoning model, with 7B parameters, and 128K context length, requiring about 14GB storage, with a 91.40 score on MATH-500.

DataLearner on WeChat

Follow DataLearner on WeChat for AI model updates and research notes.

DataLearner WeChat QR code