See key specs and per-benchmark scores for each model/mode. Scroll horizontally for all columns. 当前对比 2 个模型的评测数据与核心参数。

Claude Opus 4.6
Anthropic

Qwen3.7 Max
阿里巴巴
Each axis is a category average, normalized to a 100-point radar.
Relative edge: 指令跟随 +14.9 / Relative gap: 编程与软件工程 -5.4
Relative edge: 编程与软件工程 +5.4 / Relative gap: 指令跟随 -14.9
Method: for each model and benchmark, the chart first averages all scores in the current mode scope instead of taking the best score, then averages those benchmark scores within each category. Only benchmarks with at least two selected models scored are included; missing values are not counted as zero.
Best overall
Claude Opus 4.6 · 209.88
Best single
Claude Opus 4.6 · Text Arena (Coding) 1555.35
Modality coverage
Claude Opus 4.6 · 2 modalities
Head to head
11
Benchmarks
5
Wins
6
Losses
+0.16
Average diff
Compare benchmark results across thinking modes and tool usage.
Data sourced primarily from official releases (GitHub, Hugging Face, papers), then benchmark leaderboards, then third-party evaluators. Learn about our data methodology
Complete scores for each model/mode across selected benchmarks.
11 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.
| Benchmark | Claude Opus 4.6 | Qwen3.7 Max |
|---|---|---|
HLE 综合评估 | 53.00Extended Thinking | Tools | 53.50Thinking Enabled | Tools |
LiveBench 综合评估 | 76.33Thinking Level · High | 74.29Deep Thinking Mode |
LiveCodeBench 编程与软件工程 | 76.00Extended Thinking | 91.60Thinking Level · High |
SWE-bench Multilingual 编程与软件工程 | 72.00Extended Thinking | Tools | 78.30Thinking Enabled | Tools |
SWE-bench Verified 编程与软件工程 | 80.84Extended Thinking | Tools | 80.40Thinking Enabled | Tools |
Text Arena (Coding) 编程与软件工程 | 1555.35Standard Mode | 1540.77Standard Mode |
GPQA Diamond 科学与综合推理 | 91.31Extended Thinking | 92.40Thinking Level · High |
SimpleBench 常识推理 | 67.60Standard Mode | 70.40Standard Mode |
IF Bench 指令跟随 | 94.00Extended Thinking | 79.10Thinking Level · High |
MCP-Atlas AI Agent - 工具使用 | 76.80Thinking Level · High | Tools | 76.40Thinking Enabled | Tools |
Terminal Bench 2.0 AI Agent - 工具使用 | 65.40Extended Thinking | Tools | 69.70Thinking Enabled | Tools |
Side-by-side input/output token pricing
Licensing, MoE architecture, and multi-modality support.
| Features & specs | Claude Opus 4.6Anthropic | Qwen3.7 Max阿里巴巴 |
|---|---|---|
Core specsRelease | 2026-02-05 | 2026-05-20 |
Context length | 1000K | 1M |
Max output | 65536 | 65536 |
MoE | No | No |
LicenseCode Open Source | Closed Source | Closed Source |
Weights Open Source | Closed Source | Closed Source |
Commercial use | 不开源 | 不开源 |
Local deploymentWeight size | 0B | Not provided |
VRAM for weights | Not provided | Not provided |
Modality supportText Input/Output | / | / |
Image Input/Output | / | Not provided |
ResourcesPaper / report | Introducing Claude Opus 4.6 | Qwen3.7: The Agent Frontier |