DataLearner logo

Qwen3.8-Flash-NextvsQwen3.8-27B

Across 9 shared benchmarks, Qwen3.8-Flash-Next leads overall: Qwen3.8-Flash-Next wins 9, Qwen3.8-27B wins 0, with 0 ties and an average score difference of +4.19.

阿里巴巴
Qwen3.8-Flash-Next

阿里巴巴 · 2026-08-26 · Reasoning model

阿里巴巴
Qwen3.8-27B

阿里巴巴 · 2026-08-14 · Reasoning model

Qwen3.8-Flash-Next9 wins(100%)(0%)0 winsQwen3.8-27B

Benchmark scores

Grouped by capability, sorted by largest gap within each. 9 shared benchmarks.

Coding and Software Engineer

Qwen3.8-Flash-Next 4/4
BenchmarkQwen3.8-Flash-NextQwen3.8-27BDiff
DeepSWE58.7016 / 30极高强度思考(工具)42.2025 / 30Thinking (With Tools)+16.50
NL2Repo-Bench48.108 / 10极高强度思考(工具)42.309 / 10Thinking (With Tools)+5.80
LiveCodeBench91.903 / 127极高强度思考(无工具)90.307 / 127Thinking (No Tools)+1.60
SWE-Bench Pro - Public62.509 / 58极高强度思考(工具)61.7011 / 58Thinking (With Tools)+0.80

Multimodal Understanding

Qwen3.8-Flash-Next 2/2
BenchmarkQwen3.8-Flash-NextQwen3.8-27BDiff
MathVision95.702 / 12极高强度思考(工具)94.603 / 12Thinking (With Tools)+1.10
CharXiv RQ90.602 / 18极高强度思考(工具)90.203 / 18Thinking (With Tools)+0.40

Agent Level Benchmark

Qwen3.8-Flash-Next 1/1
BenchmarkQwen3.8-Flash-NextQwen3.8-27BDiff
Agents' Last Exam24.3012 / 13极高强度思考(工具)20.4013 / 13Thinking (With Tools)+3.90

General Evaluation

Qwen3.8-Flash-Next 1/1
BenchmarkQwen3.8-Flash-NextQwen3.8-27BDiff
GPQA Diamond91.7027 / 225极高强度思考(无工具)89.2048 / 225Thinking (No Tools)+2.50

General Knowledge

Qwen3.8-Flash-Next 1/1
BenchmarkQwen3.8-Flash-NextQwen3.8-27BDiff
HLE35.9078 / 183极高强度思考(无工具)30.8090 / 183Thinking (No Tools)+5.10

Specs

FieldQwen3.8-Flash-NextQwen3.8-27B
Publisher阿里巴巴阿里巴巴
Release date2026-08-262026-08-14
Model typeReasoning modelReasoning model
ArchitectureMoEDense
Parameters125B27B
Context length262144256K
Max output128K128K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemQwen3.8-Flash-NextQwen3.8-27B
Text inputNot public$0.5 / 1M tokens
Text outputNot public$3 / 1M tokens
Cache readNot public$0.1 / 1M tokens

One or both models have incomplete public pricing.

Summary

  • Qwen3.8-Flash-Nextleads in:Coding and Software Engineer (4/4), Multimodal Understanding (2/2), Agent Level Benchmark (1/1), General Evaluation (1/1), General Knowledge (1/1)

On average across the 9 shared benchmarks, Qwen3.8-Flash-Next scores 4.19 higher.

Largest single-benchmark gap: DeepSWE — Qwen3.8-Flash-Next 58.70 vs Qwen3.8-27B 42.20 (+16.50).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.