Qwen3.8-27BvsQwen3.5-27B
Across 4 shared benchmarks, Qwen3.8-27B leads overall: Qwen3.8-27B wins 3, Qwen3.5-27B wins 1, with 0 ties and an average score difference of +5.93.
Qwen3.8-27B
阿里巴巴 · 2026-08-14 · Reasoning model
Qwen3.5-27B
阿里巴巴 · 2026-02-25 · Reasoning model
Qwen3.8-27B3 wins(75%)(25%)1 winQwen3.5-27B
Benchmark scores
Grouped by capability, sorted by largest gap within each. 4 shared benchmarks.
AI Agent - Tool Usage
Qwen3.8-27B 1/1| Benchmark | Qwen3.8-27B | Qwen3.5-27B | Diff |
|---|---|---|---|
| OSWorld-Verified | 84.303 / 26Thinking (With Tools) | 56.2023 / 26Thinking (With Tools) | +28.10 |
Coding and Software Engineer
Qwen3.8-27B 1/1| Benchmark | Qwen3.8-27B | Qwen3.5-27B | Diff |
|---|---|---|---|
| LiveCodeBench | 90.306 / 124Thinking (No Tools) | 80.7028 / 124Thinking (With Tools) | +9.60 |
General Evaluation
Qwen3.8-27B 1/1| Benchmark | Qwen3.8-27B | Qwen3.5-27B | Diff |
|---|---|---|---|
| GPQA Diamond | 89.2047 / 226Thinking (No Tools) | 85.5084 / 226Thinking (No Tools) | +3.70 |
General Knowledge
Qwen3.5-27B 1/1| Benchmark | Qwen3.8-27B | Qwen3.5-27B | Diff |
|---|---|---|---|
| HLE | 30.8087 / 179Thinking (No Tools) | 48.5035 / 179Thinking (With Tools) | -17.70 |
Specs
| Field | Qwen3.8-27B | Qwen3.5-27B |
|---|---|---|
| Publisher | 阿里巴巴 | 阿里巴巴 |
| Release date | 2026-08-14 | 2026-02-25 |
| Model type | Reasoning model | Reasoning model |
| Architecture | Dense | Dense |
| Parameters | 27B | 27B |
| Context length | 256K | 1010K |
| Max output | 128K | 248320 |
Summary
- Qwen3.8-27Bleads in:AI Agent - Tool Usage (1/1), Coding and Software Engineer (1/1), General Evaluation (1/1)
- Qwen3.5-27Bleads in:General Knowledge (1/1)
On average across the 4 shared benchmarks, Qwen3.8-27B scores 5.93 higher.
Largest single-benchmark gap: OSWorld-Verified — Qwen3.8-27B 84.30 vs Qwen3.5-27B 56.20 (+28.10).
Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.