DataLearner logo

Qwen3.5-397B-A17B Benchmark Details

Qwen3.5-397B-A17B currently shows benchmark results led by MMLU Pro (10 / 132, score 87.80), Pinch Bench (3 / 37, score 89.10), IF Bench (4 / 30, score 76.50). 1 source link is attached for reference.

Benchmark Results

Qwen3.5-397B-A17B

Benchmark Results

Thinking
Tool usage
Internet

General Knowledge

5 evaluations
Benchmark / mode
Score
Rank/total
C-Eval
Thinking Mode
93
3 / 10
GPQA Diamond
Thinking Mode
88.40
29 / 187
MMLU Pro
Thinking Mode
87.80
10 / 132
HLE
Thinking Mode
28.70
94 / 172
HLE
Thinking ModeToolsInternet
48.30
35 / 172

Coding and Software Engineer

4 evaluations
Benchmark / mode
Score
Rank/total
LiveCodeBench
Thinking Mode
83.60
20 / 123
SWE-bench Verified
Thinking ModeTools
76.40
33 / 112
69.30
20 / 23
50.90
39 / 54

Multimodal Understanding

1 evaluations
Benchmark / mode
Score
Rank/total
MMMU
Thinking Mode
85
5 / 29

Agent Level Benchmark

1 evaluations
Benchmark / mode
Score
Rank/total
τ²-Bench
Thinking ModeTools
86.70
7 / 43

Instruction Following

1 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Thinking Mode
76.50
4 / 30

AI Agent - Information Search

2 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
Thinking ModeToolsInternet
78.60
19 / 53
BrowseComp
Thinking ModeTools
69
29 / 53

AI Agent - Tool Usage

3 evaluations
Benchmark / mode
Score
Rank/total
OSWorld-Verified
Thinking ModeTools
62.20
19 / 24
Terminal Bench 2.0
Thinking ModeTools
52.50
30 / 47
Tool Decathlon
Thinking ModeTools
38.30
7 / 9

Math and Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
AIME 2026
Thinking Mode
91.30
13 / 18
IMO-AnswerBench
Thinking Mode
80.90
17 / 21

Long Context

2 evaluations
Benchmark / mode
Score
Rank/total
AA-LCR
Thinking Mode
68.70
7 / 15
LongBench v2
Standard Mode
63.20
2 / 11

Claw-style Agent Evaluation

1 evaluations
Benchmark / mode
Score
Rank/total
Pinch Bench
Thinking ModeTools
89.10
3 / 37

Sources