DataLearner logo

DeepSeek-R1-0528-Qwen3-8B Benchmark Details

DeepSeek-R1-0528-Qwen3-8B currently shows benchmark results led by AIME2025 (129 / 215, score 63.70), LiveCodeBench (163 / 250, score 51.30), GPQA Diamond (372 / 462, score 61.20).

Benchmark Results

DeepSeek-R1-0528-Qwen3-8B

Benchmark Results

Thinking
Tool usage

General Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
HLE
Thinking Enabled
5.90
462 / 563

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Enabled
61.20
372 / 462

Coding and Software Engineer

1 evaluations
Benchmark / mode
Score
Rank/total
LiveCodeBench
Thinking Enabled
51.30
163 / 250

Math and Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
AIME2025
Thinking Enabled
63.70
129 / 215

Instruction Following

1 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Thinking Enabled
19.90
280 / 282

Agent Level Benchmark

1 evaluations
Benchmark / mode
Score
Rank/total
Terminal Bench Hard
Thinking EnabledTools
1.50
237 / 244