DataLearner logo

Qwen3.8-Flash-NextvsDeepSeek-V4-Flash

Across 9 shared benchmarks, Qwen3.8-Flash-Next leads overall: Qwen3.8-Flash-Next wins 7, DeepSeek-V4-Flash wins 2, with 0 ties and an average score difference of +12.24.

阿里巴巴
Qwen3.8-Flash-Next

阿里巴巴 · 2026-08-26 · Reasoning model

DeepSeek-AI
DeepSeek-V4-Flash

DeepSeek-AI · 2026-04-24 · Reasoning model

Qwen3.8-Flash-Next7 wins(78%)(22%)2 winsDeepSeek-V4-Flash

Benchmark scores

Grouped by capability, sorted by largest gap within each. 9 shared benchmarks.

Coding and Software Engineer

Qwen3.8-Flash-Next 4/5
BenchmarkQwen3.8-Flash-NextDeepSeek-V4-FlashDiff
LiveCodeBench91.903 / 127极高强度思考(无工具)55.2087 / 127Normal (No Tools)+36.70
SWE-Bench Pro - Public62.509 / 58极高强度思考(工具)49.1048 / 58Normal (With Tools)+13.40
SWE-bench Multilingual813 / 26极高强度思考(工具)69.7021 / 26Normal (With Tools)+11.30
NL2Repo-Bench48.108 / 10极高强度思考(工具)54.206 / 10Max (With Tools)-6.10
DeepSWE58.7016 / 30极高强度思考(工具)54.4018 / 30Max (With Tools)+4.30

Agent Level Benchmark

DeepSeek-V4-Flash 1/1
BenchmarkQwen3.8-Flash-NextDeepSeek-V4-FlashDiff
Agents' Last Exam24.3012 / 13极高强度思考(工具)25.2011 / 13Max (With Tools)-0.90

AI Agent - Tool Usage

Qwen3.8-Flash-Next 1/1
BenchmarkQwen3.8-Flash-NextDeepSeek-V4-FlashDiff
Toolathlon-Verified73.504 / 7极高强度思考(工具)70.307 / 7Max (With Tools)+3.20

General Evaluation

Qwen3.8-Flash-Next 1/1
BenchmarkQwen3.8-Flash-NextDeepSeek-V4-FlashDiff
GPQA Diamond91.7027 / 225极高强度思考(无工具)71.20150 / 225Normal (No Tools)+20.50

General Knowledge

Qwen3.8-Flash-Next 1/1
BenchmarkQwen3.8-Flash-NextDeepSeek-V4-FlashDiff
HLE35.9078 / 183极高强度思考(无工具)8.10166 / 183Normal (No Tools)+27.80

Specs

FieldQwen3.8-Flash-NextDeepSeek-V4-Flash
Publisher阿里巴巴DeepSeek-AI
Release date2026-08-262026-04-24
Model typeReasoning modelReasoning model
ArchitectureMoEMoE
Parameters125B284B
Context length2621441M
Max output128K384K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemQwen3.8-Flash-NextDeepSeek-V4-Flash
Text inputNot public$0.14 / 1M tokens
Text outputNot public$0.28 / 1M tokens
Cache readNot public$0.0028 / 1M tokens

One or both models have incomplete public pricing.

Summary

  • Qwen3.8-Flash-Nextleads in:Coding and Software Engineer (4/5), AI Agent - Tool Usage (1/1), General Evaluation (1/1), General Knowledge (1/1)
  • DeepSeek-V4-Flashleads in:Agent Level Benchmark (1/1)

On average across the 9 shared benchmarks, Qwen3.8-Flash-Next scores 12.24 higher.

Largest single-benchmark gap: LiveCodeBench — Qwen3.8-Flash-Next 91.90 vs DeepSeek-V4-Flash 55.20 (+36.70).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.