DataLearner logo

Qwen3.7 Max Benchmark Details

Qwen3.7 Max currently shows benchmark results led by MMLU Pro (4 / 134, score 89.60), LiveCodeBench (5 / 128, score 91.60), GPQA Diamond (27 / 274, score 92.40). This page also compares it with 3 competitor models and 3 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

Qwen3.7 Max

Benchmark Results

Thinking
Tool usage

General Knowledge

4 evaluations
Benchmark / mode
Score
Rank/total
MMLU Pro
Thinking Level · Max
89.60
4 / 134
LiveBench
Deep Thinking Mode
74.29
21 / 115
HLE
Thinking ModeTools
53.50
24 / 197
HLE
Thinking Level · Max
41.40
71 / 197

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Level · Max
92.40
27 / 274

Coding and Software Engineer

6 evaluations
Benchmark / mode
Score
Rank/total
Text Arena (Coding)
Standard Mode
1540.77
11 / 35
LiveCodeBench
Thinking Level · Max
91.60
5 / 128
SWE-bench Verified
Thinking ModeTools
80.40
13 / 116
SWE-bench Multilingual
Thinking ModeTools
78.30
7 / 29
SWE-Bench Pro - Public
Thinking ModeTools
60.60
15 / 62
SciCode
Thinking Level · Max
53.50
8 / 16

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Standard Mode
70.40
14 / 92

Instruction Following

1 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Thinking Level · Max
79.10
5 / 36

AI Agent - Tool Usage

2 evaluations
Benchmark / mode
Score
Rank/total
MCP-Atlas
Thinking ModeTools
76.40
18 / 41
Terminal Bench 2.0
Thinking ModeTools
69.70
5 / 48

Text Embedding

1 evaluations
Benchmark / mode
Score
Rank/total
Context Arena
Standard Mode
56
79 / 126

Math and Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
IMO-AnswerBench
Thinking Level · Max
90
3 / 24

Competitor Comparison

Benchmark scores for Qwen3.7 Max compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkQwen3.7 MaxCurrentClaude Opus 4.6Kimi K2.6GLM 5.1
HLE
Accuracy
综合评估
53.50Thinking Enabled | Tools
53.00Extended Thinking | Tools
54.00Thinking Enabled | Tools
52.30Thinking Enabled | Tools
LiveBench
Accuracy
综合评估
74.29Deep Thinking Mode
76.33Thinking Level · High
72.17Thinking Enabled
70.18Standard Mode
GPQA Diamond
Accuracy
科学与综合推理
92.40Thinking Level · High
91.31Extended Thinking
90.50Thinking Enabled
86.20Thinking Enabled
LiveCodeBench
Pass @K
编程与软件工程
91.60Thinking Level · High
76.00Extended Thinking
89.60Thinking Enabled
--
SWE-bench Multilingual
Accuracy
编程与软件工程
78.30Thinking Enabled | Tools
72.00Extended Thinking | Tools
76.70Thinking Enabled | Tools
--
SWE-Bench Pro - Public
Accuracy
编程与软件工程
60.60Thinking Enabled | Tools
--
58.60Thinking Enabled | Tools
58.40Thinking Enabled | Tools
SWE-bench Verified
Accuracy
编程与软件工程
80.40Thinking Enabled | Tools
80.84Extended Thinking | Tools
80.20Thinking Enabled | Tools
--
Text Arena (Coding)
Arena Score
编程与软件工程
1540.77Standard Mode
1555.35Standard Mode
--
1534.00Standard Mode
SimpleBench
Score (AVG@5)
常识推理
70.40Standard Mode
67.60Standard Mode
--
55.10Standard Mode
IF Bench
Accuracy
指令跟随
79.10Thinking Level · High
94.00Extended Thinking
--
--
MCP-Atlas
Pass rate / claim coverage
AI Agent - 工具使用
76.40Thinking Enabled | Tools
76.80Thinking Level · High | Tools
69.40Thinking Enabled | Tools
75.60Standard Mode | Tools
Terminal Bench 2.0
Accuracy
AI Agent - 工具使用
69.70Thinking Enabled | Tools
65.40Extended Thinking | Tools
66.70Thinking Enabled | Tools
63.50Thinking Enabled | Tools
2 additional benchmarks remain in the chart above.

Standard API Pricing: Qwen3.7 Max vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

Qwen3.7 Max
Supplier: 阿里巴巴
Standard input: ¥12 / 1M tokens
Standard output: ¥36 / 1M tokens
Claude Opus 4.6
Supplier: Anthropic
Standard input: $5 / 1M tokens
Standard output: $25 / 1M tokens
Kimi K2.6
Supplier: Facebook AI研究实验室
Standard input: $0.95 / 1M tokens
Standard output: $4 / 1M tokens
GLM 5.1
Supplier: 智谱AI
Standard input: $1.4 / 1M tokens
Standard output: $4.4 / 1M tokens
ModelSupplierStandard inputStandard outputBase price applies to
Qwen3.7 Max
阿里巴巴¥12 / 1M tokens¥36 / 1M tokens
Claude Opus 4.6
Anthropic$5 / 1M tokens$25 / 1M tokens
Kimi K2.6
Facebook AI研究实验室$0.95 / 1M tokens$4 / 1M tokens
GLM 5.1
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens

Version History

How each version of the Qwen3.7 Max series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkQwen3.7 MaxCurrentQwen3.6-Max-PreviewQwen3-Max-Thinking
HLE
Accuracy
综合评估
53.50Thinking Enabled | Tools
50.20Thinking Enabled | Tools
49.80Thinking Enabled | Tools
MMLU Pro
Accuracy
综合评估
89.60Thinking Level · High
88.50Thinking Level · High
85.70Thinking Enabled
GPQA Diamond
Accuracy
科学与综合推理
92.40Thinking Level · High
90.40Thinking Level · High
87.40Thinking Enabled
LiveCodeBench
Pass @K
编程与软件工程
91.60Thinking Level · High
87.10Thinking Level · High
85.90Thinking Enabled
SWE-bench Multilingual
Accuracy
编程与软件工程
78.30Thinking Enabled | Tools
73.80Thinking Enabled | Tools
--
SWE-Bench Pro - Public
Accuracy
编程与软件工程
60.60Thinking Enabled | Tools
57.30Deep Thinking Mode | Tools
--
SWE-bench Verified
Accuracy
编程与软件工程
80.40Thinking Enabled | Tools
78.80Thinking Enabled | Tools
75.30Thinking Enabled
SimpleBench
Score (AVG@5)
常识推理
70.40Standard Mode
63.00Standard Mode
--
IF Bench
Accuracy
指令跟随
79.10Thinking Level · High
74.20Thinking Level · High
70.90Thinking Enabled | Tools
Terminal Bench 2.0
Accuracy
AI Agent - 工具使用
69.70Thinking Enabled | Tools
65.40Deep Thinking Mode | Tools
--
Context Arena
Accuracy (8 needles, 4K-128K context)
文本向量检索
56.01Standard Mode
88.14Thinking Enabled
--
IMO-AnswerBench
Accuracy
数学推理
90.00Thinking Level · High
83.80Thinking Level · High
83.90Thinking Enabled

Single-Benchmark Version Trend

Viewing: HLE · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Qwen3.7 Max Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

Qwen3.7 Max
Supplier: 阿里巴巴
Standard input: ¥12 / 1M tokens
Standard output: ¥36 / 1M tokens
Qwen3.6-Max-Preview
Supplier: 阿里巴巴
Standard input: $1.3 / 1M tokens
Standard output: $7.8 / 1M tokens
Base price applies to <= 128
Qwen3-Max-Thinking
Supplier: 阿里巴巴
Standard input: $1.2 / 1M tokens
Standard output: $6 / 1M tokens
Base price applies to <= 32000
ModelSupplierStandard inputStandard outputBase price applies to
Qwen3.7 Max
阿里巴巴¥12 / 1M tokens¥36 / 1M tokens
Qwen3.6-Max-Preview
阿里巴巴$1.3 / 1M tokens$7.8 / 1M tokens<= 128
Qwen3-Max-Thinking
阿里巴巴$1.2 / 1M tokens$6 / 1M tokens<= 32000