DataLearner logo

Qwen2.5-7B Benchmark Details

Qwen2.5-7B currently shows benchmark results led by MBPP (27 / 96, score 74.90), GSM8K (21 / 71, score 85.40), ARC (2 / 4, score 63.70).

Benchmark Results

Qwen2.5-7B

Benchmark Results

Thinking
Tool usage

General Knowledge

2 evaluations
Benchmark / mode
Score
Rank/total
MMLU
Standard Mode
74.20
67 / 124
MMLU-Pro
Standard Mode
45
147 / 176

Math and Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
GSM8K
Standard Mode
85.40
21 / 71
MATH
Standard Mode
49.80
34 / 42

Coding and Software Engineer

2 evaluations
Benchmark / mode
Score
Rank/total
MBPP
Standard Mode
74.90
27 / 96
HumanEval
Standard Mode
57.90
76 / 140

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
ARC
Standard Mode
63.70
2 / 4

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
36.40
439 / 463

Claw-style Agent Evaluation

1 evaluations
Benchmark / mode
Score
Rank/total
Pinch Bench
Thinking EnabledTools
40.30
38 / 38