DataLearner logo

Qwen3.5-9B Benchmark Details

Qwen3.5-9B currently shows benchmark results led by C-Eval (8 / 48, score 88.20), MMLU Pro (52 / 133, score 82.50), GPQA Diamond (130 / 270, score 81.70).

Benchmark Results

Qwen3.5-9B

Benchmark Results

Thinking

General Knowledge

2 evaluations
Benchmark / mode
Score
Rank/total
C-Eval
Thinking Enabled
88.20
8 / 48
MMLU Pro
Thinking Enabled
82.50
52 / 133

Other

2 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
78.91
151 / 270
GPQA Diamond
Thinking Enabled
81.70
130 / 270

Coding and Software Engineer

1 evaluations
Benchmark / mode
Score
Rank/total
LiveCodeBench
Thinking Enabled
65.60
66 / 127

Instruction Following

1 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Thinking Enabled
64.50
24 / 34

Text Embedding

2 evaluations
Benchmark / mode
Score
Rank/total
Context Arena
Standard Mode
22.60
121 / 126
Context Arena
Thinking Enabled
43.28
97 / 126

Long Context

2 evaluations
Benchmark / mode
Score
Rank/total
AA-LCR
Thinking Enabled
63
24 / 27
LongBench v2
Thinking Enabled
55.20
11 / 13