Qwen 3.6 Plus PreviewvsGLM 5.1
Across 9 shared benchmarks, GLM 5.1 leads overall: Qwen 3.6 Plus Preview wins 2, GLM 5.1 wins 5, with 2 ties and an average score difference of +1.51.
Qwen 3.6 Plus Preview
阿里巴巴 · 2026-03-31 · Chat model
GLM 5.1
智谱AI · 2026-03-27 · Reasoning model
Qwen 3.6 Plus Preview2 wins(22%)Ties2(56%)5 winsGLM 5.1
Benchmark scores
Grouped by capability, sorted by largest gap within each. 9 shared benchmarks.
AI Agent - Tool Usage
GLM 5.1 3/3| Benchmark | Qwen 3.6 Plus Preview | GLM 5.1 | Diff |
|---|---|---|---|
| Terminal Bench 2.0 | 61.6016 / 48Thinking (With Tools) | 63.5013 / 48Thinking (With Tools) | -1.90 |
| Tool Decathlon | 39.807 / 10Thinking (With Tools) | 40.706 / 10Thinking (With Tools) | -0.90 |
| Terminal-Bench 2.1 | 61.40106 / 191Thinking (With Tools) | 61.80104 / 191Thinking (With Tools) | -0.40 |
General Knowledge
GLM 5.1 2/2| Benchmark | Qwen 3.6 Plus Preview | GLM 5.1 | Diff |
|---|---|---|---|
| CritPt | 2.90115 / 200Thinking (No Tools) | 4.60102 / 200Thinking (No Tools) | -1.70 |
| LiveBench | 68.9142 / 117Normal (No Tools) | 70.1837 / 117Normal (No Tools) | -1.27 |
Math and Reasoning
Even 2/2| Benchmark | Qwen 3.6 Plus Preview | GLM 5.1 | Diff |
|---|---|---|---|
| AIME 2026 | 95.3012 / 29Thinking (No Tools) | 95.3012 / 29Thinking (No Tools) | — |
| IMO-AnswerBench | 83.8014 / 24Thinking (No Tools) | 83.8014 / 24Thinking (No Tools) | — |
Agent Level Benchmark
Qwen 3.6 Plus Preview 1/1| Benchmark | Qwen 3.6 Plus Preview | GLM 5.1 | Diff |
|---|---|---|---|
| τ³-Banking | 20.8088 / 164Thinking (With Tools) | 13.60115 / 164Thinking (With Tools) | +7.20 |
Claw-style Agent Evaluation
Qwen 3.6 Plus Preview 1/1| Benchmark | Qwen 3.6 Plus Preview | GLM 5.1 | Diff |
|---|---|---|---|
| PinchBench v2 | 72.5023 / 45Reported best (effort unspecified) | 59.9533 / 45Reported best (effort unspecified) | +12.55 |
Specs
| Field | Qwen 3.6 Plus Preview | GLM 5.1 |
|---|---|---|
| Publisher | 阿里巴巴 | 智谱AI |
| Release date | 2026-03-31 | 2026-03-27 |
| Model type | Chat model | Reasoning model |
| Architecture | Dense | MoE |
| Parameters | Not available | 754B |
| Context length | 1M | 200K |
| Max output | 64K | 125K |
API pricing
Prices use DataLearner records when available; missing fields are not inferred.
| Item | Qwen 3.6 Plus Preview | GLM 5.1 |
|---|---|---|
| Text input | $0.5 / 1M tokens | $1.4 / 1M tokens |
| Text output | $3 / 1M tokens | $4.4 / 1M tokens |
| Cache read | $0.05 / 1M tokens | $4.4 / 1M tokens |
| Cache write | $0.625 / 1M tokens | $0.26 / 1M tokens |
Summary
- Qwen 3.6 Plus Previewleads in:Agent Level Benchmark (1/1), Claw-style Agent Evaluation (1/1)
- GLM 5.1leads in:AI Agent - Tool Usage (3/3), General Knowledge (2/2)
- Tied in:Math and Reasoning
On average across the 9 shared benchmarks, Qwen 3.6 Plus Preview scores 1.51 higher.
Largest single-benchmark gap: PinchBench v2 — Qwen 3.6 Plus Preview 72.50 vs GLM 5.1 59.95 (+12.55).
Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.