DataLearner logo

Qwen3.8-Flash-Next Benchmark Details

Qwen3.8-Flash-Next currently shows benchmark results led by LiveCodeBench (3 / 127, score 91.90), IF Bench (3 / 35, score 81.30), CharXiv RQ (2 / 19, score 90.60). This page also compares it with 2 competitor models and 1 predecessor or same-series models, including performance and pricing views when available. 3 source links are attached for reference.

Benchmark Results

Qwen3.8-Flash-Next

Benchmark Results

Thinking
Tool usage

General Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
HLE
Extra-High
35.90
84 / 190

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Extra-High
91.70
32 / 271

Coding and Software Engineer

5 evaluations
Benchmark / mode
Score
Rank/total
LiveCodeBench
Extra-High
91.90
3 / 127
SWE-bench Multilingual
Extra-HighTools
81
4 / 28
SWE-Bench Pro - Public
Extra-HighTools
62.50
10 / 60
DeepSWE
Extra-HighTools
58.70
21 / 35
NL2Repo-Bench
Extra-HighTools
48.10
10 / 12

Instruction Following

1 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Extra-High
81.30
3 / 35

AI Agent - Tool Usage

5 evaluations
Benchmark / mode
Score
Rank/total
AndroidWorld
Extra-HighTools
84.50
1 / 1
Toolathlon-Verified
Extra-HighTools
73.50
6 / 10
ClawEval-MM
Extra-HighTools
60.40
2 / 2
OSWorld 2.0
Extra-HighTools
52.30
8 / 9
RecreationBench
Extra-HighTools
49.90
1 / 1

Agent Level Benchmark

3 evaluations
Benchmark / mode
Score
Rank/total
CoWorkBench
Extra-HighTools
73.90
2 / 2
Job Bench
Extra-HighTools
55.70
4 / 5
Agents' Last Exam
Extra-HighTools
24.30
13 / 15

Multimodal Understanding

8 evaluations
Benchmark / mode
Score
Rank/total
MathVision
Extra-High
90.60
7 / 12
MathVision
Extra-HighTools
95.70
2 / 12
CharXiv RQ
Extra-High
84.60
13 / 19
CharXiv RQ
Extra-HighTools
90.60
2 / 19
RealWorldQA
Extra-High
88.50
1 / 1
LVBench
Extra-High
76.60
4 / 4
ERQA
Extra-High
72.30
2 / 2
Vision2Web
Extra-HighTools
64
1 / 1

Competitor Comparison

Benchmark scores for Qwen3.8-Flash-Next compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

11 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkQwen3.8-Flash-NextCurrentQwen3.8-27BDeepSeek-V4-Flash
HLE
综合评估
35.90Thinking Level · Extra High
30.80Thinking Enabled
45.10Thinking Level · Extra High | Tools
GPQA Diamond
科学与综合推理
91.70Thinking Level · Extra High
89.20Thinking Enabled
88.10Thinking Level · High
DeepSWE
编程与软件工程
58.70Thinking Level · Extra High | Tools
42.20Thinking Enabled | Tools
54.40Thinking Level · High | Tools
LiveCodeBench
编程与软件工程
91.90Thinking Level · Extra High
90.30Thinking Enabled
91.60Thinking Level · High
NL2Repo-Bench
编程与软件工程
48.10Thinking Level · Extra High | Tools
42.30Thinking Enabled | Tools
54.20Thinking Level · High | Tools
SWE-bench Multilingual
编程与软件工程
81.00Thinking Level · Extra High | Tools
--
73.30Thinking Level · Extra High | Tools
SWE-Bench Pro - Public
编程与软件工程
62.50Thinking Level · Extra High | Tools
61.70Thinking Enabled | Tools
52.60Thinking Level · Extra High | Tools
Toolathlon-Verified
AI Agent - 工具使用
73.50Thinking Level · Extra High | Tools
--
70.30Thinking Level · High | Tools
Agents' Last Exam
Agent能力评测
24.30Thinking Level · Extra High | Tools
20.40Thinking Enabled | Tools
25.20Thinking Level · High | Tools
CharXiv RQ
多模态理解
90.60Thinking Level · Extra High | Tools
90.20Thinking Enabled | Tools
--
MathVision
多模态理解
95.70Thinking Level · Extra High | Tools
94.60Thinking Enabled | Tools
--

Standard API Pricing: Qwen3.8-Flash-Next vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

Qwen3.8-Flash-Next
Supplier: 阿里巴巴
Standard input: ¥1 / 1M tokens
Standard output: ¥3 / 1M tokens
DeepSeek-V4-Flash
Supplier: DeepSeek-AI
Standard input: $0.14 / 1M tokens
Standard output: $0.28 / 1M tokens
ModelSupplierStandard inputStandard outputBase price applies to
Qwen3.8-Flash-Next
阿里巴巴¥1 / 1M tokens¥3 / 1M tokens
DeepSeek-V4-Flash
DeepSeek-AI$0.14 / 1M tokens$0.28 / 1M tokens

Version History

How each version of the Qwen3.8-Flash-Next series stacks up on benchmark tests

Qwen3.8-Flash-NextQwen3.7-Plus
Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

1 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkQwen3.8-Flash-NextCurrentQwen3.7-Plus
GPQA Diamond
科学与综合推理
91.70Thinking Level · Extra High
81.82Standard Mode

Single-Benchmark Version Trend

Viewing: GPQA Diamond · 科学与综合推理

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Qwen3.8-Flash-Next Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · CNY / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Qwen3.7-Plus: Base price applies to <= 256000
ModelSupplierStandard inputStandard outputBase price applies to
Qwen3.8-Flash-Next
阿里巴巴¥1 / 1M tokens¥3 / 1M tokens
Qwen3.7-Plus
阿里巴巴¥2 / 1M tokens¥8 / 1M tokens<= 256000

Sources