Qwen3.8-Flash-NextvsDeepSeek-V4-Flash
Across 9 shared benchmarks, Qwen3.8-Flash-Next leads overall: Qwen3.8-Flash-Next wins 7, DeepSeek-V4-Flash wins 2, with 0 ties and an average score difference of +12.24.
阿里巴巴 · 2026-08-26 · Reasoning model
DeepSeek-AI · 2026-04-24 · Reasoning model
Benchmark scores
Grouped by capability, sorted by largest gap within each. 9 shared benchmarks.
Coding and Software Engineer
Qwen3.8-Flash-Next 4/5| Benchmark | Qwen3.8-Flash-Next | DeepSeek-V4-Flash | Diff |
|---|---|---|---|
| LiveCodeBench | 91.903 / 127极高强度思考(无工具) | 55.2087 / 127Normal (No Tools) | +36.70 |
| SWE-Bench Pro - Public | 62.509 / 58极高强度思考(工具) | 49.1048 / 58Normal (With Tools) | +13.40 |
| SWE-bench Multilingual | 813 / 26极高强度思考(工具) | 69.7021 / 26Normal (With Tools) | +11.30 |
| NL2Repo-Bench | 48.108 / 10极高强度思考(工具) | 54.206 / 10Max (With Tools) | -6.10 |
| DeepSWE | 58.7016 / 30极高强度思考(工具) | 54.4018 / 30Max (With Tools) | +4.30 |
Agent Level Benchmark
DeepSeek-V4-Flash 1/1| Benchmark | Qwen3.8-Flash-Next | DeepSeek-V4-Flash | Diff |
|---|---|---|---|
| Agents' Last Exam | 24.3012 / 13极高强度思考(工具) | 25.2011 / 13Max (With Tools) | -0.90 |
AI Agent - Tool Usage
Qwen3.8-Flash-Next 1/1| Benchmark | Qwen3.8-Flash-Next | DeepSeek-V4-Flash | Diff |
|---|---|---|---|
| Toolathlon-Verified | 73.504 / 7极高强度思考(工具) | 70.307 / 7Max (With Tools) | +3.20 |
General Evaluation
Qwen3.8-Flash-Next 1/1| Benchmark | Qwen3.8-Flash-Next | DeepSeek-V4-Flash | Diff |
|---|---|---|---|
| GPQA Diamond | 91.7027 / 225极高强度思考(无工具) | 71.20150 / 225Normal (No Tools) | +20.50 |
General Knowledge
Qwen3.8-Flash-Next 1/1| Benchmark | Qwen3.8-Flash-Next | DeepSeek-V4-Flash | Diff |
|---|---|---|---|
| HLE | 35.9078 / 183极高强度思考(无工具) | 8.10166 / 183Normal (No Tools) | +27.80 |
Specs
| Field | Qwen3.8-Flash-Next | DeepSeek-V4-Flash |
|---|---|---|
| Publisher | 阿里巴巴 | DeepSeek-AI |
| Release date | 2026-08-26 | 2026-04-24 |
| Model type | Reasoning model | Reasoning model |
| Architecture | MoE | MoE |
| Parameters | 125B | 284B |
| Context length | 262144 | 1M |
| Max output | 128K | 384K |
API pricing
Prices use DataLearner records when available; missing fields are not inferred.
| Item | Qwen3.8-Flash-Next | DeepSeek-V4-Flash |
|---|---|---|
| Text input | Not public | $0.14 / 1M tokens |
| Text output | Not public | $0.28 / 1M tokens |
| Cache read | Not public | $0.0028 / 1M tokens |
One or both models have incomplete public pricing.
Summary
- Qwen3.8-Flash-Nextleads in:Coding and Software Engineer (4/5), AI Agent - Tool Usage (1/1), General Evaluation (1/1), General Knowledge (1/1)
- DeepSeek-V4-Flashleads in:Agent Level Benchmark (1/1)
On average across the 9 shared benchmarks, Qwen3.8-Flash-Next scores 12.24 higher.
Largest single-benchmark gap: LiveCodeBench — Qwen3.8-Flash-Next 91.90 vs DeepSeek-V4-Flash 55.20 (+36.70).
Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.