DataLearner logo

Qwen3.5-27B Benchmark Details

Qwen3.5-27B currently shows benchmark results led by Pinch Bench (2 / 38, score 90), C-Eval (6 / 48, score 90.50), MMLU Pro (20 / 133, score 86.10). This page also compares it with 1 competitor models and 2 predecessor or same-series models, including performance and pricing views when available. 1 source link is attached for reference.

Benchmark Results

Qwen3.5-27B

Benchmark Results

Thinking
Tool usage
Internet

General Knowledge

4 evaluations
Benchmark / mode
Score
Rank/total
C-Eval
Thinking Mode
90.50
6 / 48
MMLU Pro
Thinking Mode
86.10
20 / 133
HLE
Thinking Mode
24.30
117 / 185
HLE
Thinking ModeTools
48.50
37 / 185

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Mode
85.50
95 / 270

Coding and Software Engineer

3 evaluations
Benchmark / mode
Score
Rank/total
CodeForces
Thinking Mode
1899
16 / 20
LiveCodeBench
Thinking ModeTools
80.70
30 / 127
SWE-bench Verified
Thinking Mode
72.40
54 / 114

Multimodal Understanding

2 evaluations
Benchmark / mode
Score
Rank/total
MMMU
Thinking Mode
82.30
8 / 28
SimpleVQA
Thinking Mode
56
3 / 3

Agent Level Benchmark

1 evaluations
Benchmark / mode
Score
Rank/total
τ²-Bench
Thinking ModeTools
79
17 / 43

Instruction Following

1 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Thinking Mode
76.50
7 / 34

AI Agent - Information Search

2 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
Thinking ModeTools
61
34 / 54
BrowseComp
Thinking ModeToolsInternet
61
34 / 54

AI Agent - Tool Usage

2 evaluations
Benchmark / mode
Score
Rank/total
OSWorld-Verified
Thinking ModeTools
56.20
23 / 26
Terminal Bench 2.0
Thinking ModeTools
41.60
44 / 48

Text Embedding

2 evaluations
Benchmark / mode
Score
Rank/total
Context Arena
Standard Mode
36.82
102 / 126
Context Arena
Thinking Mode
73.34
56 / 126

Long Context

2 evaluations
Benchmark / mode
Score
Rank/total
AA-LCR
Thinking Mode
66.10
17 / 27
LongBench v2
Standard Mode
60.60
8 / 13

Claw-style Agent Evaluation

2 evaluations
Benchmark / mode
Score
Rank/total
Pinch Bench
Thinking ModeTools
90
2 / 38
Claw Bench
Thinking ModeTools
75.20
26 / 29

Competitor Comparison

Benchmark scores for Qwen3.5-27B compared against top models in its class

Qwen3.5-27BGemma 4 31B
Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

6 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkQwen3.5-27BCurrentGemma 4 31B
HLE
综合评估
48.50Thinking Enabled | Tools
26.50Thinking Enabled | Tools
MMLU Pro
综合评估
86.10Thinking Enabled
85.20Thinking Enabled
GPQA Diamond
科学与综合推理
85.50Thinking Enabled
84.30Thinking Enabled
CodeForces
编程与软件工程
1899.00Thinking Enabled
2150.00Thinking Enabled
LiveCodeBench
编程与软件工程
80.70Thinking Enabled | Tools
80.00Thinking Enabled
τ²-Bench
Agent能力评测
79.00Thinking Enabled | Tools
76.90Thinking Enabled | Tools

Standard API Pricing: Qwen3.5-27B vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

Comparable standard text pricing is not available for these models.

Version History

How each version of the Qwen3.5-27B series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

5 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkQwen3.5-27BCurrentQwen3-32BQwen2.5-32B
C-Eval
综合评估
90.50Thinking Enabled
87.30Thinking Enabled
--
MMLU Pro
综合评估
86.10Thinking Enabled
--
69.23Standard Mode
GPQA Diamond
科学与综合推理
85.50Thinking Enabled
68.40Thinking Enabled
--
CodeForces
编程与软件工程
1899.00Thinking Enabled
1977.00Thinking Enabled
--
LiveCodeBench
编程与软件工程
80.70Thinking Enabled | Tools
65.70Thinking Enabled
51.20Standard Mode

Single-Benchmark Version Trend

Viewing: C-Eval · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Qwen3.5-27B Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

Qwen3-32B
Supplier: 阿里巴巴
Standard input: ¥0.0012 / 1K tokens
Standard output: ¥0.0048 / 1K tokens
Qwen2.5-32B
Supplier: 阿里巴巴
Standard input: ¥0.002 / 1K tokens
Standard output: ¥0.006 / 1K tokens
ModelSupplierStandard inputStandard outputBase price applies to
Qwen3-32B
阿里巴巴¥0.0012 / 1K tokens¥0.0048 / 1K tokens
Qwen2.5-32B
阿里巴巴¥0.002 / 1K tokens¥0.006 / 1K tokens

Sources