DataLearner logo

Qwen3.8-Max Benchmark Details

Qwen3.8-Max currently shows benchmark results led by Context Arena (3 / 126, score 96.92), LongBench v2 (1 / 13, score 66.30), HLE (15 / 189, score 56.20). This page also compares it with 3 competitor models and 3 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

Qwen3.8-Max

Benchmark Results

Thinking
Tool usage

General Knowledge

2 evaluations
Benchmark / mode
Score
Rank/total
HLE
Extra-High
43.60
54 / 189
HLE
Extra-HighTools
56.20
15 / 189

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Extra-High
92.60
24 / 270

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1842
10 / 99

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Extra-High
62.50
22 / 92

Coding and Software Engineer

5 evaluations
Benchmark / mode
Score
Rank/total
FrontierSWE
Extra-HighTools
73.50
4 / 4
SWE-Bench Pro - Public
Extra-HighTools
67.70
5 / 60
DeepSWE
Extra-HighTools
56.60
21 / 34
NL2Repo-Bench
Extra-HighTools
55.90
7 / 12
MLS Bench
Extra-HighTools
41
3 / 5

Text Embedding

3 evaluations
Benchmark / mode
Score
Rank/total
92.79
9 / 126
Context Arena
Thinking Level · Medium
95.15
5 / 126
Context Arena
Extra-High
96.92
3 / 126

AI Agent - Tool Usage

2 evaluations
Benchmark / mode
Score
Rank/total
Terminal-Bench 2.1
Extra-HighTools
86.60
10 / 49
Toolathlon-Verified
Extra-HighTools
72.50
9 / 10

Long Context

1 evaluations
Benchmark / mode
Score
Rank/total
LongBench v2
Extra-High
66.30
1 / 13

Agent Level Benchmark

1 evaluations
Benchmark / mode
Score
Rank/total
Agents' Last Exam
Extra-HighTools
27
7 / 14

Productivity Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
AutomationBench
Extra-HighTools
27.30
12 / 16

Math and Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
FrontierMath v2
Extra-High
74.74
11 / 58
46.34
11 / 40

Competitor Comparison

Benchmark scores for Qwen3.8-Max compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkQwen3.8-MaxCurrentClaude Opus 5GPT-5.6 SolKimi K3
HLE
Accuracy
综合评估
56.20Thinking Level · Extra High | Tools
64.70Thinking Level · High | Tools
49.50Thinking Level · High
56.00Thinking Level · High | Tools
GPQA Diamond
Accuracy
科学与综合推理
92.60Thinking Level · Extra High
93.88Thinking Level · High
93.50Thinking Level · High
93.50Thinking Level · High
Creative Writing
Elo、大模型评判两两对战
写作和创作
1842.00Standard Mode
2116.10Standard Mode
1964.10Standard Mode
2070.80Standard Mode
SimpleBench
Score (AVG@5)
常识推理
62.50Thinking Level · Extra High
80.60Thinking Level · High
64.80Thinking Level · Extra High
60.70Thinking Level · High
DeepSWE
Pass@1 (DeepSWE v1.1)
编程与软件工程
56.60Thinking Level · Extra High | Tools
68.80Thinking Level · High | Tools
72.70Thinking Level · Extra High | Tools
67.50Thinking Level · High | Tools
FrontierSWE
Dominance score
编程与软件工程
73.50Thinking Level · Extra High | Tools
--
--
81.20Thinking Level · High | Tools
MLS Bench
Score
编程与软件工程
41.00Thinking Level · Extra High | Tools
--
--
48.30Thinking Level · High | Tools
SWE-Bench Pro - Public
Accuracy
编程与软件工程
67.70Thinking Level · Extra High | Tools
79.20Thinking Level · High | Tools
64.60Thinking Level · Extra High | Tools
--
Context Arena
Accuracy (8 needles, 4K-128K context)
文本向量检索
96.92Thinking Level · Extra High
97.72Thinking Level · High
97.63Thinking Level · High
71.75Thinking Level · High
Terminal-Bench 2.1
Accuracy
AI Agent - 工具使用
86.60Thinking Level · Extra High | Tools
--
88.80Thinking Level · High
88.30Thinking Level · High | Tools
Toolathlon-Verified
Score
AI Agent - 工具使用
72.50Thinking Level · Extra High | Tools
--
--
76.50Thinking Level · High | Tools
Agents' Last Exam
Score
Agent能力评测
27.00Thinking Level · Extra High | Tools
--
52.70Thinking Level · Extra High | Tools
28.30Thinking Level · High | Tools
3 additional benchmarks remain in the chart above.

Standard API Pricing: Qwen3.8-Max vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

Qwen3.8-Max
Supplier: 阿里巴巴
Standard input: ¥12 / 1M tokens
Standard output: ¥36 / 1M tokens
Claude Opus 5
Supplier: Anthropic
Standard input: $5 / 1M tokens
Standard output: $25 / 1M tokens
GPT-5.6 Sol
Supplier: OpenAI
Standard input: $4 / 1M tokens
Standard output: $20 / 1M tokens
Kimi K3
Supplier: Moonshot AI
Standard input: ¥20 / 1M tokens
Standard output: ¥100 / 1M tokens
ModelSupplierStandard inputStandard outputBase price applies to
Qwen3.8-Max
阿里巴巴¥12 / 1M tokens¥36 / 1M tokens
Claude Opus 5
Anthropic$5 / 1M tokens$25 / 1M tokens
GPT-5.6 Sol
OpenAI$4 / 1M tokens$20 / 1M tokens
Kimi K3
Moonshot AI¥20 / 1M tokens¥100 / 1M tokens

Version History

How each version of the Qwen3.8-Max series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

5 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkQwen3.8-MaxCurrentQwen3.7 MaxQwen3.6-Max-Preview
HLE
Accuracy
综合评估
56.20Thinking Level · Extra High | Tools
53.50Thinking Enabled | Tools
50.20Thinking Enabled | Tools
GPQA Diamond
Accuracy
科学与综合推理
92.60Thinking Level · Extra High
92.40Thinking Level · High
90.40Thinking Level · High
SimpleBench
Score (AVG@5)
常识推理
62.50Thinking Level · Extra High
70.40Standard Mode
63.00Standard Mode
SWE-Bench Pro - Public
Accuracy
编程与软件工程
67.70Thinking Level · Extra High | Tools
60.60Thinking Enabled | Tools
57.30Deep Thinking Mode | Tools
Context Arena
Accuracy (8 needles, 4K-128K context)
文本向量检索
96.92Thinking Level · Extra High
56.01Standard Mode
88.14Thinking Enabled

Single-Benchmark Version Trend

Viewing: HLE · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Qwen3.8-Max Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

Qwen3.8-Max
Supplier: 阿里巴巴
Standard input: ¥12 / 1M tokens
Standard output: ¥36 / 1M tokens
Qwen3.7 Max
Supplier: 阿里巴巴
Standard input: ¥12 / 1M tokens
Standard output: ¥36 / 1M tokens
Qwen3.6-Max-Preview
Supplier: 阿里巴巴
Standard input: $1.3 / 1M tokens
Standard output: $7.8 / 1M tokens
Base price applies to <= 128
ModelSupplierStandard inputStandard outputBase price applies to
Qwen3.8-Max
阿里巴巴¥12 / 1M tokens¥36 / 1M tokens
Qwen3.7 Max
阿里巴巴¥12 / 1M tokens¥36 / 1M tokens
Qwen3.6-Max-Preview
阿里巴巴$1.3 / 1M tokens$7.8 / 1M tokens<= 128