Qwen3.8-Flash-NextvsQwen3.7-Plus
Across 8 shared benchmarks, Qwen3.8-Flash-Next leads overall: Qwen3.8-Flash-Next wins 7, Qwen3.7-Plus wins 1, with 0 ties and an average score difference of +11.11.
Qwen3.8-Flash-Next
阿里巴巴 · 2026-08-26 · Reasoning model
Qwen3.7-Plus
阿里巴巴 · 2026-05-31 · Reasoning model
Qwen3.8-Flash-Next7 wins(88%)(13%)1 winQwen3.7-Plus
Benchmark scores
Grouped by capability, sorted by largest gap within each. 8 shared benchmarks.
AI Agent - Tool Usage
Qwen3.8-Flash-Next 2/2| Benchmark | Qwen3.8-Flash-Next | Qwen3.7-Plus | Diff |
|---|---|---|---|
| Terminal-Bench 2.1 | 86.1025 / 192Thinking (With Tools) | 61108 / 192Thinking (With Tools) | +25.10 |
| Terminal-Bench 4.0 | 25.3022 / 86Thinking (With Tools) | 174 / 86Thinking (With Tools) | +24.30 |
General Knowledge
Qwen3.8-Flash-Next 2/2| Benchmark | Qwen3.8-Flash-Next | Qwen3.7-Plus | Diff |
|---|---|---|---|
| HLE | 38144 / 563Thinking (No Tools) · Text only | 35.60166 / 563Thinking (No Tools) · Text only | +2.40 |
| CritPt | 11.1067 / 200Thinking (No Tools) | 9.1076 / 200Thinking (No Tools) | +2 |
Multimodal Understanding
Even 2/2| Benchmark | Qwen3.8-Flash-Next | Qwen3.7-Plus | Diff |
|---|---|---|---|
| GDP.pdf | 15.6056 / 118Thinking (No Tools) | 12.2070 / 118Thinking (No Tools) | +3.40 |
| MMMU-Pro | 79.8046 / 227Thinking (No Tools) | 80.5037 / 227Thinking (No Tools) | -0.70 |
Agent Level Benchmark
Qwen3.8-Flash-Next 1/1| Benchmark | Qwen3.8-Flash-Next | Qwen3.7-Plus | Diff |
|---|---|---|---|
| τ³-Banking | 45.4016 / 164Thinking (With Tools) | 17.5098 / 164Thinking (With Tools) | +27.90 |
Coding and Software Engineer
Qwen3.8-Flash-Next 1/1| Benchmark | Qwen3.8-Flash-Next | Qwen3.7-Plus | Diff |
|---|---|---|---|
| SciCode | 50.6065 / 130Thinking (No Tools) | 46.1084 / 130Thinking (No Tools) | +4.50 |
Specs
| Field | Qwen3.8-Flash-Next | Qwen3.7-Plus |
|---|---|---|
| Publisher | 阿里巴巴 | 阿里巴巴 |
| Release date | 2026-08-26 | 2026-05-31 |
| Model type | Reasoning model | Reasoning model |
| Architecture | MoE | Dense |
| Parameters | 125B | Not available |
| Context length | 262144 | 1M |
| Max output | 128K | 64K |
API pricing
Prices use DataLearner records when available; missing fields are not inferred.
| Item | Qwen3.8-Flash-Next | Qwen3.7-Plus |
|---|---|---|
| Text input | ¥1 / 1M tokens | ¥2 / 1M tokens |
| Text output | ¥3 / 1M tokens | ¥8 / 1M tokens |
Summary
- Qwen3.8-Flash-Nextleads in:AI Agent - Tool Usage (2/2), General Knowledge (2/2), Agent Level Benchmark (1/1), Coding and Software Engineer (1/1)
- Tied in:Multimodal Understanding
On average across the 8 shared benchmarks, Qwen3.8-Flash-Next scores 11.11 higher.
Largest single-benchmark gap: τ³-Banking — Qwen3.8-Flash-Next 45.40 vs Qwen3.7-Plus 17.50 (+27.90).
Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.