DataLearner logo

Qwen3-Max-Thinking Benchmark Details

Qwen3-Max-Thinking currently shows benchmark results led by C-Eval (1 / 10, score 93.70), LiveCodeBench (14 / 123, score 85.90), MMLU Pro (22 / 132, score 85.70).

Benchmark Results

Qwen3-Max-Thinking

Benchmark Results

Thinking
Tool usage

General Knowledge

5 evaluations
Benchmark / mode
Score
Rank/total
93.70
1 / 10
87.40
38 / 187
85.70
22 / 132
49.80
30 / 172
30.20
88 / 172

Coding and Software Engineer

2 evaluations
Benchmark / mode
Score
Rank/total
85.90
14 / 123
75.30
37 / 112

Agent Level Benchmark

1 evaluations
Benchmark / mode
Score
Rank/total
82.10
11 / 43

Instruction Following

1 evaluations
Benchmark / mode
Score
Rank/total
70.90
12 / 30

Math and Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
83.90
11 / 21

Claw-style Agent Evaluation

1 evaluations
Benchmark / mode
Score
Rank/total
Pinch Bench
Thinking EnabledTools
80.30
23 / 37