DataLearner logo

Qwen3.5-27BvsQwen3-32B

Across 3 shared benchmarks, Qwen3.5-27B leads overall: Qwen3.5-27B wins 3, Qwen3-32B wins 0, with 0 ties and an average score difference of +18.27.

阿里巴巴
Qwen3.5-27B

阿里巴巴 · 2026-02-25 · Reasoning model

阿里巴巴
Qwen3-32B

阿里巴巴 · 2025-04-28 · Reasoning model

Qwen3.5-27B3 wins(100%)(0%)0 winsQwen3-32B

Benchmark scores

Grouped by capability, sorted by largest gap within each. 3 shared benchmarks.

General Evaluation

Qwen3.5-27B 1/1
BenchmarkQwen3.5-27BQwen3-32BDiff
GPQA Diamond84.20174 / 462Normal (No Tools)54.60401 / 462Normal (No Tools)+29.60

General Knowledge

Qwen3.5-27B 1/1
BenchmarkQwen3.5-27BQwen3-32BDiff
HLE13.90348 / 563Normal (No Tools) · Text only4.10519 / 563Normal (No Tools) · Text only+9.80

Instruction Following

Qwen3.5-27B 1/1
BenchmarkQwen3.5-27BQwen3-32BDiff
IF Bench46.90171 / 282Normal (No Tools)31.50262 / 282Normal (No Tools)+15.40

Specs

FieldQwen3.5-27BQwen3-32B
Publisher阿里巴巴阿里巴巴
Release date2026-02-252025-04-28
Model typeReasoning modelReasoning model
ArchitectureDenseDense
Parameters27B32B
Context length1010K128K
Max output24832016K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemQwen3.5-27BQwen3-32B
Text inputNot public¥0.0012 / 1K tokens
Text outputNot public¥0.0048 / 1K tokens

One or both models have incomplete public pricing.

Summary

  • Qwen3.5-27Bleads in:General Evaluation (1/1), General Knowledge (1/1), Instruction Following (1/1)

On average across the 3 shared benchmarks, Qwen3.5-27B scores 18.27 higher.

Largest single-benchmark gap: GPQA Diamond — Qwen3.5-27B 84.20 vs Qwen3-32B 54.60 (+29.60).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.