DataLearner logo

Qwen3.6-Max-Preview Benchmark Details

Qwen3.6-Max-Preview currently shows benchmark results led by MMLU Pro (5 / 134, score 88.50), LiveCodeBench (13 / 128, score 87.10), Context Arena (19 / 126, score 88.14). This page also compares it with 3 competitor models and 2 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

Qwen3.6-Max-Preview

Benchmark Results

Thinking
Tool usage

General Knowledge

3 evaluations
Benchmark / mode
Score
Rank/total
MMLU Pro
Thinking Level · Max
88.50
5 / 134
HLE
Thinking ModeTools
50.20
34 / 191
HLE
Thinking Level · Max
28.80
106 / 191

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Level · Max
90.40
44 / 273

Coding and Software Engineer

5 evaluations
Benchmark / mode
Score
Rank/total
LiveCodeBench
Thinking Level · Max
87.10
13 / 128
SWE-bench Verified
Thinking ModeTools
78.80
21 / 116
SWE-bench Multilingual
Thinking ModeTools
73.80
13 / 29
SWE-Bench Pro - Public
Thinking ModeTools
56.60
25 / 62
SWE-Bench Pro - Public
Deep Thinking ModeTools
57.30
23 / 62

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Standard Mode
63
21 / 92

Instruction Following

1 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Thinking Level · Max
74.20
11 / 36

AI Agent - Tool Usage

2 evaluations
Benchmark / mode
Score
Rank/total
Terminal Bench 2.0
Thinking ModeTools
61.60
16 / 48
Terminal Bench 2.0
Deep Thinking ModeTools
65.40
11 / 48

Text Embedding

2 evaluations
Benchmark / mode
Score
Rank/total
Context Arena
Standard Mode
70.62
61 / 126
Context Arena
Thinking Mode
88.14
19 / 126

Math and Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
IMO-AnswerBench
Thinking Level · Max
83.80
14 / 24

Productivity Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA
Deep Thinking Mode
51
12 / 21

Competitor Comparison

Benchmark scores for Qwen3.6-Max-Preview compared against top models in its class

Qwen3.6-Max-PreviewGLM 5.1Kimi K2.6Opus 4.7
Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

10 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkQwen3.6-Max-PreviewCurrentGLM 5.1Kimi K2.6Opus 4.7
HLE
Accuracy
综合评估
50.20Thinking Enabled | Tools
52.30Thinking Enabled | Tools
54.00Thinking Enabled | Tools
54.70Extended Thinking | Tools
GPQA Diamond
Accuracy
科学与综合推理
90.40Thinking Level · High
86.20Thinking Enabled
90.50Thinking Enabled
94.20Extended Thinking
LiveCodeBench
Pass @K
编程与软件工程
87.10Thinking Level · High
--
89.60Thinking Enabled
--
SWE-bench Multilingual
Accuracy
编程与软件工程
73.80Thinking Enabled | Tools
--
76.70Thinking Enabled | Tools
--
SWE-Bench Pro - Public
Accuracy
编程与软件工程
57.30Deep Thinking Mode | Tools
58.40Thinking Enabled | Tools
58.60Thinking Enabled | Tools
64.30Extended Thinking | Tools
SWE-bench Verified
Accuracy
编程与软件工程
78.80Thinking Enabled | Tools
--
80.20Thinking Enabled | Tools
87.60Extended Thinking | Tools
SimpleBench
Score (AVG@5)
常识推理
63.00Standard Mode
55.10Standard Mode
--
61.70Standard Mode
Terminal Bench 2.0
Accuracy
AI Agent - 工具使用
65.40Deep Thinking Mode | Tools
63.50Thinking Enabled | Tools
66.70Thinking Enabled | Tools
69.40Extended Thinking | Tools
Context Arena
Accuracy (8 needles, 4K-128K context)
文本向量检索
88.14Thinking Enabled
62.05Thinking Enabled
64.63Thinking Enabled
46.70Thinking Level · Low
IMO-AnswerBench
Accuracy
数学推理
83.80Thinking Level · High
83.80Thinking Enabled
86.00Thinking Enabled
--

Standard API Pricing: Qwen3.6-Max-Preview vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Qwen3.6-Max-Preview: Base price applies to <= 128
ModelSupplierStandard inputStandard outputBase price applies to
Qwen3.6-Max-Preview
阿里巴巴$1.3 / 1M tokens$7.8 / 1M tokens<= 128
GLM 5.1
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens
Kimi K2.6
Facebook AI研究实验室$0.95 / 1M tokens$4 / 1M tokens
Opus 4.7
Anthropic$5 / 1M tokens$25 / 1M tokens

Version History

How each version of the Qwen3.6-Max-Preview series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

7 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkQwen3.6-Max-PreviewCurrentQwen3-Max-Thinking
HLE
Accuracy
综合评估
50.20Thinking Enabled | Tools
49.80Thinking Enabled | Tools
MMLU Pro
Accuracy
综合评估
88.50Thinking Level · High
85.70Thinking Enabled
GPQA Diamond
Accuracy
科学与综合推理
90.40Thinking Level · High
87.40Thinking Enabled
LiveCodeBench
Pass @K
编程与软件工程
87.10Thinking Level · High
85.90Thinking Enabled
SWE-bench Verified
Accuracy
编程与软件工程
78.80Thinking Enabled | Tools
75.30Thinking Enabled
IF Bench
Accuracy
指令跟随
74.20Thinking Level · High
70.90Thinking Enabled | Tools
IMO-AnswerBench
Accuracy
数学推理
83.80Thinking Level · High
83.90Thinking Enabled

Single-Benchmark Version Trend

Viewing: HLE · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Qwen3.6-Max-Preview Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Qwen3.6-Max-Preview: Base price applies to <= 128
Qwen3-Max-Thinking: Base price applies to <= 32000
ModelSupplierStandard inputStandard outputBase price applies to
Qwen3.6-Max-Preview
阿里巴巴$1.3 / 1M tokens$7.8 / 1M tokens<= 128
Qwen3-Max-Thinking
阿里巴巴$1.2 / 1M tokens$6 / 1M tokens<= 32000