DataLearner logo

Qwen3 Max (Preview) Benchmark Details

Qwen3 Max (Preview) currently shows benchmark results led by MMLU Pro (40 / 176, score 84), τ²-Bench - Telecom (96 / 264, score 84.20), AIME2025 (97 / 215, score 80.60).

Benchmark Results

Qwen3 Max (Preview)

Benchmark Results

Thinking
Tool usage

General Knowledge

2 evaluations
Benchmark / mode
Score
Rank/total
MMLU Pro
Standard Mode
84
40 / 176
HLE
Standard Mode
11.10
376 / 563

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
76
272 / 462

Coding and Software Engineer

2 evaluations
Benchmark / mode
Score
Rank/total
SWE-bench Verified
Standard Mode
69.60
63 / 116
LiveCodeBench
Standard Mode
57.50
139 / 250

Math and Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
AIME2025
Standard Mode
80.60
97 / 215

AI Agent - Tool Usage

1 evaluations
Benchmark / mode
Score
Rank/total
Terminal-Bench
Standard ModeTools
19
32 / 35

Agent Level Benchmark

3 evaluations
Benchmark / mode
Score
Rank/total
τ²-Bench - Telecom
Thinking ModeTools
84.20
96 / 264
τ²-Bench
Standard ModeTools
74
23 / 44
τ²-Bench
Thinking ModeTools
72
25 / 44