DataLearner logo

Qwen3.8-Flash-NextvsQwen3.7-Plus

Across 8 shared benchmarks, Qwen3.8-Flash-Next leads overall: Qwen3.8-Flash-Next wins 7, Qwen3.7-Plus wins 1, with 0 ties and an average score difference of +11.11.

阿里巴巴
Qwen3.8-Flash-Next

阿里巴巴 · 2026-08-26 · Reasoning model

阿里巴巴
Qwen3.7-Plus

阿里巴巴 · 2026-05-31 · Reasoning model

Qwen3.8-Flash-Next7 wins(88%)(13%)1 winQwen3.7-Plus

Benchmark scores

Grouped by capability, sorted by largest gap within each. 8 shared benchmarks.

AI Agent - Tool Usage

Qwen3.8-Flash-Next 2/2
BenchmarkQwen3.8-Flash-NextQwen3.7-PlusDiff
Terminal-Bench 2.186.1025 / 192Thinking (With Tools)61108 / 192Thinking (With Tools)+25.10
Terminal-Bench 4.025.3022 / 86Thinking (With Tools)174 / 86Thinking (With Tools)+24.30

General Knowledge

Qwen3.8-Flash-Next 2/2
BenchmarkQwen3.8-Flash-NextQwen3.7-PlusDiff
HLE38144 / 563Thinking (No Tools) · Text only35.60166 / 563Thinking (No Tools) · Text only+2.40
CritPt11.1067 / 200Thinking (No Tools)9.1076 / 200Thinking (No Tools)+2

Multimodal Understanding

Even 2/2
BenchmarkQwen3.8-Flash-NextQwen3.7-PlusDiff
GDP.pdf15.6056 / 118Thinking (No Tools)12.2070 / 118Thinking (No Tools)+3.40
MMMU-Pro79.8046 / 227Thinking (No Tools)80.5037 / 227Thinking (No Tools)-0.70

Agent Level Benchmark

Qwen3.8-Flash-Next 1/1
BenchmarkQwen3.8-Flash-NextQwen3.7-PlusDiff
τ³-Banking45.4016 / 164Thinking (With Tools)17.5098 / 164Thinking (With Tools)+27.90

Coding and Software Engineer

Qwen3.8-Flash-Next 1/1
BenchmarkQwen3.8-Flash-NextQwen3.7-PlusDiff
SciCode50.6065 / 130Thinking (No Tools)46.1084 / 130Thinking (No Tools)+4.50

Specs

FieldQwen3.8-Flash-NextQwen3.7-Plus
Publisher阿里巴巴阿里巴巴
Release date2026-08-262026-05-31
Model typeReasoning modelReasoning model
ArchitectureMoEDense
Parameters125BNot available
Context length2621441M
Max output128K64K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemQwen3.8-Flash-NextQwen3.7-Plus
Text input¥1 / 1M tokens¥2 / 1M tokens
Text output¥3 / 1M tokens¥8 / 1M tokens

Summary

  • Qwen3.8-Flash-Nextleads in:AI Agent - Tool Usage (2/2), General Knowledge (2/2), Agent Level Benchmark (1/1), Coding and Software Engineer (1/1)
  • Tied in:Multimodal Understanding

On average across the 8 shared benchmarks, Qwen3.8-Flash-Next scores 11.11 higher.

Largest single-benchmark gap: τ³-Banking — Qwen3.8-Flash-Next 45.40 vs Qwen3.7-Plus 17.50 (+27.90).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.