DataLearner logo

Qwen3.6-27BvsQwen3.5-27B

Across 13 shared benchmarks, Qwen3.6-27B leads overall: Qwen3.6-27B wins 10, Qwen3.5-27B wins 3, with 0 ties and an average score difference of +2.65.

阿里巴巴
Qwen3.6-27B

阿里巴巴 · 2026-04-22 · Reasoning model

阿里巴巴
Qwen3.5-27B

阿里巴巴 · 2026-02-25 · Reasoning model

Qwen3.6-27B10 wins(77%)(23%)3 winsQwen3.5-27B

Benchmark scores

Grouped by capability, sorted by largest gap within each. 13 shared benchmarks.

General Knowledge

Qwen3.6-27B 4/4
BenchmarkQwen3.6-27BQwen3.5-27BDiff
HLE15.10333 / 563Normal (No Tools) · Text only13.90348 / 563Normal (No Tools) · Text only+1.20
C-Eval91.405 / 48Thinking (No Tools)90.506 / 48Thinking (No Tools)+0.90
CritPt0.90158 / 200Normal (No Tools)0.30182 / 200Normal (No Tools)+0.60
MMLU Pro86.2019 / 176Thinking (No Tools)86.1021 / 176Thinking (No Tools)+0.10

Agent Level Benchmark

Even 2/2
BenchmarkQwen3.6-27BQwen3.5-27BDiff
Terminal Bench Hard21.20145 / 244Normal (With Tools)31.80100 / 244Normal (With Tools)-10.60
τ²-Bench - Telecom93.6046 / 264Normal (With Tools)87.1074 / 264Normal (With Tools)+6.50

AI Agent - Tool Usage

Qwen3.6-27B 1/1
BenchmarkQwen3.6-27BQwen3.5-27BDiff
Terminal Bench 2.059.3020 / 48Thinking (With Tools)41.6044 / 48Thinking (With Tools)+17.70

Claw-style Agent Evaluation

Qwen3.5-27B 1/1
BenchmarkQwen3.6-27BQwen3.5-27BDiff
Claw Bench72.4027 / 29Thinking (With Tools)75.2026 / 29Thinking (With Tools)-2.80

General Evaluation

Qwen3.6-27B 1/1
BenchmarkQwen3.6-27BQwen3.5-27BDiff
GPQA Diamond84.85159 / 462Normal (No Tools)84.20174 / 462Normal (No Tools)+0.65

Instruction Following

Qwen3.5-27B 1/1
BenchmarkQwen3.6-27BQwen3.5-27BDiff
IF Bench45.70177 / 282Normal (No Tools)46.90171 / 282Normal (No Tools)-1.20

Long Context

Qwen3.6-27B 1/1
BenchmarkQwen3.6-27BQwen3.5-27BDiff
AA-LCR66.70118 / 170Normal (No Tools)64122 / 170Normal (No Tools)+2.70

Multimodal Understanding

Qwen3.6-27B 1/1
BenchmarkQwen3.6-27BQwen3.5-27BDiff
MMMU-Pro71.70118 / 227Normal (No Tools)70130 / 227Normal (No Tools)+1.70

Text Embedding

Qwen3.6-27B 1/1
BenchmarkQwen3.6-27BQwen3.5-27BDiff
Context Arena53.7882 / 126Normal (No Tools)36.82102 / 126Normal (No Tools)+16.96

Specs

FieldQwen3.6-27BQwen3.5-27B
Publisher阿里巴巴阿里巴巴
Release date2026-04-222026-02-25
Model typeReasoning modelReasoning model
ArchitectureDenseDense
Parameters27B27B
Context length128K1010K
Max output16K248320

Summary

  • Qwen3.6-27Bleads in:General Knowledge (4/4), AI Agent - Tool Usage (1/1), General Evaluation (1/1), Long Context (1/1), Multimodal Understanding (1/1), Text Embedding (1/1)
  • Qwen3.5-27Bleads in:Claw-style Agent Evaluation (1/1), Instruction Following (1/1)
  • Tied in:Agent Level Benchmark

On average across the 13 shared benchmarks, Qwen3.6-27B scores 2.65 higher.

Largest single-benchmark gap: Terminal Bench 2.0 — Qwen3.6-27B 59.30 vs Qwen3.5-27B 41.60 (+17.70).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.