DataLearner logo

Qwen2.5-Max Benchmark Details

Qwen2.5-Max currently shows benchmark results led by GSM8K (9 / 70, score 94.50), MMLU (22 / 124, score 87.90), MBPP (19 / 96, score 80.60).

Benchmark Results

Qwen2.5-Max

Benchmark Results

Thinking

General Knowledge

3 evaluations
Benchmark / mode
Score
Rank/total
MMLU
Standard Mode
87.90
22 / 124
MMLU Pro
Standard Mode
76.10
82 / 176
HLE
Standard Mode
3.80
533 / 563

Math and Reasoning

3 evaluations
Benchmark / mode
Score
Rank/total
GSM8K
Standard Mode
94.50
9 / 70
MATH
Standard Mode
68.50
24 / 42
FrontierMath
Standard Mode
1
52 / 60

Coding and Software Engineer

3 evaluations
Benchmark / mode
Score
Rank/total
MBPP
Standard Mode
80.60
19 / 96
HumanEval
Standard Mode
73.20
51 / 140
LiveCodeBench
Standard Mode
35.90
194 / 250

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
58.70
384 / 462

Agent Level Benchmark

1 evaluations
Benchmark / mode
Score
Rank/total
Aider-Polyglot
Standard Mode
21.80
48 / 59
Qwen2.5-Max Benchmark Results & Rankings | DataLearnerAI