DataLearner logo

Qwen3.6-27B Benchmark Details

Qwen3.6-27B currently shows benchmark results led by C-Eval (5 / 48, score 91.40), MMLU Pro (19 / 176, score 86.20), LiveCodeBench (35 / 250, score 83.90). This page also compares it with 3 competitor models and 3 predecessor or same-series models, including performance and pricing views when available. 1 source link is attached for reference.

Benchmark Results

Qwen3.6-27B

Benchmark Results

Thinking
Tool usage

General Knowledge

8 evaluations
Benchmark / mode
Score
Rank/total
C-Eval
Thinking Mode
91.40
5 / 48
MMLU Pro
Thinking Mode
86.20
19 / 176
LiveBench
Standard Mode
64.03
54 / 117
HLE
Standard Mode
15.10
333 / 563
HLE
Thinking Mode
24
259 / 563
HLE
Thinking Mode
23.10
265 / 563
CritPt
Standard Mode
0.90
158 / 200
CritPt
Thinking Mode
1.10
147 / 200

Other

2 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
84.85
159 / 462
GPQA Diamond
Thinking Mode
87.80
123 / 462

Coding and Software Engineer

5 evaluations
Benchmark / mode
Score
Rank/total
LiveCodeBench
Thinking Mode
83.90
35 / 250
SWE-bench Verified
Thinking ModeTools
77.20
28 / 116
SWE-bench Multilingual
Thinking ModeTools
71.30
21 / 30
SWE-Bench Pro - Public
Thinking ModeTools
53.50
39 / 62
SciCode
Thinking Mode
42.80
96 / 130

Agent Level Benchmark

6 evaluations
Benchmark / mode
Score
Rank/total
τ²-Bench - Telecom
Standard ModeTools
93.60
46 / 264
τ²-Bench - Telecom
Thinking ModeTools
94.20
38 / 264
Terminal Bench Hard
Standard ModeTools
21.20
145 / 244
Terminal Bench Hard
Thinking ModeTools
34.80
75 / 244
τ³-Banking
Standard ModeTools
9.30
134 / 164
τ³-Banking
Thinking ModeTools
16.70
100 / 164

Instruction Following

2 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Standard Mode
45.70
177 / 282
IF Bench
Thinking Mode
67.60
88 / 282

AI Agent - Tool Usage

3 evaluations
Benchmark / mode
Score
Rank/total
Terminal-Bench 2.1
Standard ModeTools
51.30
129 / 192
Terminal-Bench 2.1
Thinking ModeTools
60.70
110 / 192
Terminal Bench 2.0
Thinking ModeTools
59.30
20 / 48

Text Embedding

2 evaluations
Benchmark / mode
Score
Rank/total
Context Arena
Standard Mode
53.78
82 / 126
Context Arena
Thinking Mode
82.17
32 / 126

Math and Reasoning

3 evaluations
Benchmark / mode
Score
Rank/total
AIME 2026
Thinking Mode
94.10
15 / 29
IMO-AnswerBench
Thinking Mode
80.80
20 / 24
FrontierMath v2
Standard Mode
34.04
41 / 58

Long Context

1 evaluations
Benchmark / mode
Score
Rank/total
AA-LCR
Standard Mode
66.70
118 / 170

Claw-style Agent Evaluation

1 evaluations
Benchmark / mode
Score
Rank/total
Claw Bench
Thinking ModeTools
72.40
27 / 29

Productivity Knowledge

2 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA v2
Standard ModeTools
1043
87 / 105
Harvey Lab-AA
Thinking ModeTools
82.30
27 / 43

Multimodal Understanding

3 evaluations
Benchmark / mode
Score
Rank/total
MMMU-Pro
Standard Mode
71.70
118 / 227
MMMU-Pro
Thinking Mode
74.60
91 / 227
GDP.pdf
Thinking Mode
11
76 / 118

Competitor Comparison

Benchmark scores for Qwen3.6-27B compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkQwen3.6-27BCurrentGemini 3.0 FlashHaiku 4.5GPT-5.4 mini
CritPt
Score
综合评估
1.10Thinking Enabled
8.60Thinking Enabled
--
10.00Thinking Level · Extra High
HLE
Accuracy
综合评估
24.00Thinking Enabled
43.50Thinking Enabled | Tools
10.40Thinking Enabled
41.50Thinking Level · Extra High | Tools
LiveBench
Accuracy
综合评估
64.03Standard Mode
72.40Thinking Level · High
61.3264K
66.37Deep Thinking Mode
MMLU Pro
Accuracy
综合评估
86.20Thinking Enabled
--
80.00Extended Thinking
--
GPQA Diamond
Accuracy
科学与综合推理
87.80Thinking Enabled
90.40Thinking Enabled
73.30Extended Thinking
87.50Thinking Level · Extra High
LiveCodeBench
Pass @K
编程与软件工程
83.90Thinking Enabled
90.80Thinking Enabled
62.00Extended Thinking
--
SciCode
Score
编程与软件工程
42.80Thinking Enabled
--
42.20Thinking Enabled
52.10Thinking Level · Extra High
SWE-Bench Pro - Public
Accuracy
编程与软件工程
53.50Thinking Enabled | Tools
49.60Thinking Level · High | Tools
39.45Extended Thinking | Tools
54.40Thinking Level · Extra High | Tools
SWE-bench Verified
Accuracy
编程与软件工程
77.20Thinking Enabled | Tools
68.70Thinking Enabled
73.30128K | Tools
--
Terminal Bench Hard
Accuracy
Agent能力评测
34.80Thinking Enabled | Tools
38.60Thinking Enabled | Tools
27.30Standard Mode | Tools
52.30Thinking Level · Extra High | Tools
τ²-Bench - Telecom
Accuracy
Agent能力评测
94.20Thinking Enabled | Tools
91.23Thinking Level · High | Tools
54.70Thinking Enabled | Tools
83.30Thinking Level · Extra High | Tools
τ³-Banking
Score
Agent能力评测
16.70Thinking Enabled | Tools
27.32Thinking Level · High | Tools
9.30Thinking Enabled | Tools
25.60Thinking Level · Extra High | Tools
11 additional benchmarks remain in the chart above.

Standard API Pricing: Qwen3.6-27B vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Gemini 3.0 Flash
Google Deep Mind$0.5 / 1M tokens$3 / 1M tokens
Haiku 4.5
Anthropic$1 / 1M tokens$5 / 1M tokens
GPT-5.4 mini
OpenAI$0.75 / 1M tokens$4.5 / 1M tokens

Version History

How each version of the Qwen3.6-27B series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkQwen3.6-27BCurrentQwen3.5-27BQwen3-32BQwen2.5-32B
C-Eval
Accuracy
综合评估
91.40Thinking Enabled
90.50Thinking Enabled
87.30Thinking Enabled
--
CritPt
Score
综合评估
1.10Thinking Enabled
0.90Thinking Enabled
0.30Thinking Enabled
--
HLE
Accuracy
综合评估
24.00Thinking Enabled
48.50Thinking Enabled | Tools
7.40Thinking Enabled
--
LiveBench
Accuracy
综合评估
64.03Standard Mode
--
43.56Thinking Enabled
--
MMLU Pro
Accuracy
综合评估
86.20Thinking Enabled
86.10Thinking Enabled
--
69.23Standard Mode
GPQA Diamond
Accuracy
科学与综合推理
87.80Thinking Enabled
85.50Thinking Enabled
68.40Thinking Enabled
--
LiveCodeBench
Pass @K
编程与软件工程
83.90Thinking Enabled
80.70Thinking Enabled | Tools
65.70Thinking Enabled
51.20Standard Mode
SciCode
Score
编程与软件工程
42.80Thinking Enabled
--
36.00Thinking Enabled
--
SWE-bench Verified
Accuracy
编程与软件工程
77.20Thinking Enabled | Tools
72.40Thinking Enabled
--
--
Terminal Bench Hard
Accuracy
Agent能力评测
34.80Thinking Enabled | Tools
32.60Thinking Enabled | Tools
3.00Thinking Enabled | Tools
--
τ²-Bench - Telecom
Accuracy
Agent能力评测
94.20Thinking Enabled | Tools
93.90Thinking Enabled | Tools
29.80Thinking Enabled | Tools
--
τ³-Banking
Score
Agent能力评测
16.70Thinking Enabled | Tools
--
5.40Thinking Enabled | Tools
--
8 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: C-Eval · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Qwen3.6-27B Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

Qwen3-32B
Supplier: 阿里巴巴
Standard input: ¥0.0012 / 1K tokens
Standard output: ¥0.0048 / 1K tokens
Qwen2.5-32B
Supplier: 阿里巴巴
Standard input: ¥0.002 / 1K tokens
Standard output: ¥0.006 / 1K tokens
ModelSupplierStandard inputStandard outputBase price applies to
Qwen3-32B
阿里巴巴¥0.0012 / 1K tokens¥0.0048 / 1K tokens
Qwen2.5-32B
阿里巴巴¥0.002 / 1K tokens¥0.006 / 1K tokens

Sources