DataLearner logo

Qwen2.5-3B Benchmark Details

Qwen2.5-3B currently shows benchmark results led by GSM8K (24 / 70, score 79.10), MBPP (30 / 70, score 57.10), HumanEval (58 / 101, score 42.10).

Benchmark Results

Qwen2.5-3B

Benchmark Results

Thinking

General Knowledge

3 evaluations
Benchmark / mode
Score
Rank/total
MMLU
Standard Mode
65.60
84 / 124
BBH
Standard Mode
56.30
17 / 21
MMLU Pro
Standard Mode
34.60
130 / 133

Math and Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
GSM8K
Standard Mode
79.10
24 / 70
MATH
Standard Mode
42.60
37 / 42

Coding and Software Engineer

2 evaluations
Benchmark / mode
Score
Rank/total
MBPP
Standard Mode
57.10
30 / 70
HumanEval
Standard Mode
42.10
58 / 101

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
24.30
221 / 224