DataLearner logo

Qwen3.8-Flash-Next Benchmark Details

Qwen3.8-Flash-Next currently shows benchmark results led by LiveCodeBench (3 / 127, score 91.90), IF Bench (3 / 34, score 81.30), CharXiv RQ (2 / 18, score 90.60). This page also compares it with 2 competitor models and 1 predecessor or same-series models, including performance and pricing views when available. 3 source links are attached for reference.

Benchmark Results

Qwen3.8-Flash-Next

Benchmark Results

Thinking
Tool usage

General Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
HLE
Extra-High
35.90
78 / 183

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Extra-High
91.70
27 / 225

Coding and Software Engineer

5 evaluations
Benchmark / mode
Score
Rank/total
LiveCodeBench
Extra-High
91.90
3 / 127
SWE-bench Multilingual
Extra-HighTools
81
3 / 26
SWE-Bench Pro - Public
Extra-HighTools
62.50
9 / 58
DeepSWE
Extra-HighTools
58.70
16 / 30
NL2Repo-Bench
Extra-HighTools
48.10
8 / 10

Instruction Following

1 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Extra-High
81.30
3 / 34

AI Agent - Tool Usage

5 evaluations
Benchmark / mode
Score
Rank/total
AndroidWorld
Extra-HighTools
84.50
1 / 1
Toolathlon-Verified
Extra-HighTools
73.50
4 / 7
ClawEval-MM
Extra-HighTools
60.40
1 / 1
OSWorld 2.0
Extra-HighTools
52.30
4 / 5
RecreationBench
Extra-HighTools
49.90
1 / 1

Agent Level Benchmark

3 evaluations
Benchmark / mode
Score
Rank/total
CoWorkBench
Extra-HighTools
73.90
1 / 1
Job Bench
Extra-HighTools
55.70
1 / 2
Agents' Last Exam
Extra-HighTools
24.30
12 / 13

Multimodal Understanding

8 evaluations
Benchmark / mode
Score
Rank/total
MathVision
Extra-High
90.60
7 / 12
MathVision
Extra-HighTools
95.70
2 / 12
CharXiv RQ
Extra-High
84.60
12 / 18
CharXiv RQ
Extra-HighTools
90.60
2 / 18
RealWorldQA
Extra-High
88.50
1 / 1
LVBench
Extra-High
76.60
2 / 2
ERQA
Extra-High
72.30
1 / 1
Vision2Web
Extra-HighTools
64
1 / 1

Competitor Comparison

Benchmark scores for Qwen3.8-Flash-Next compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

11 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkQwen3.8-Flash-NextCurrentQwen3.8-27BDeepSeek-V4-Flash
HLE
综合评估
35.90Thinking Level · Extra High
30.80Thinking Enabled
45.10Thinking Level · Extra High | Tools
GPQA Diamond
科学与综合推理
91.70Thinking Level · Extra High
89.20Thinking Enabled
88.10Thinking Level · High
DeepSWE
编程与软件工程
58.70Thinking Level · Extra High | Tools
42.20Thinking Enabled | Tools
54.40Thinking Level · High | Tools
LiveCodeBench
编程与软件工程
91.90Thinking Level · Extra High
90.30Thinking Enabled
91.60Thinking Level · High
NL2Repo-Bench
编程与软件工程
48.10Thinking Level · Extra High | Tools
42.30Thinking Enabled | Tools
54.20Thinking Level · High | Tools
SWE-bench Multilingual
编程与软件工程
81.00Thinking Level · Extra High | Tools
--
73.30Thinking Level · Extra High | Tools
SWE-Bench Pro - Public
编程与软件工程
62.50Thinking Level · Extra High | Tools
61.70Thinking Enabled | Tools
52.60Thinking Level · Extra High | Tools
Toolathlon-Verified
AI Agent - 工具使用
73.50Thinking Level · Extra High | Tools
--
70.30Thinking Level · High | Tools
Agents' Last Exam
Agent能力评测
24.30Thinking Level · Extra High | Tools
20.40Thinking Enabled | Tools
25.20Thinking Level · High | Tools
CharXiv RQ
多模态理解
90.60Thinking Level · Extra High | Tools
90.20Thinking Enabled | Tools
--
MathVision
多模态理解
95.70Thinking Level · Extra High | Tools
94.60Thinking Enabled | Tools
--

Standard API Pricing: Qwen3.8-Flash-Next vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
DeepSeek-V4-Flash
DeepSeek-AI$0.14 / 1M tokens$0.28 / 1M tokens

Version History

How each version of the Qwen3.8-Flash-Next series stacks up on benchmark tests

Qwen3.8-Flash-NextQwen3.7-Plus
No benchmark data matches the selected filters.

Standard API Pricing Across the Qwen3.8-Flash-Next Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · CNY / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Qwen3.7-Plus: Base price applies to <= 256000
ModelSupplierStandard inputStandard outputBase price applies to
Qwen3.7-Plus
阿里巴巴¥2 / 1M tokens¥8 / 1M tokens<= 256000

Sources