DataLearner logo

Qwen3.7 MaxvsClaude Opus 4.6

Across 11 shared benchmarks, Qwen3.7 Max leads overall: Qwen3.7 Max wins 6, Claude Opus 4.6 wins 5, with 0 ties and an average score difference of -0.16.

阿里巴巴
Qwen3.7 Max

阿里巴巴 · 2026-05-20 · Reasoning model

Anthropic
Claude Opus 4.6

Anthropic · 2026-02-05 · Reasoning model

Qwen3.7 Max6 wins(55%)(45%)5 winsClaude Opus 4.6

Benchmark scores

Grouped by capability, sorted by largest gap within each. 11 shared benchmarks.

Coding and Software Engineer

Even 4/4
BenchmarkQwen3.7 MaxClaude Opus 4.6Diff
LiveCodeBench91.604 / 126最高(无工具)7640 / 126Extended (no tools)+15.60
Text Arena (Coding)1,54111 / 35Normal (No Tools)1,5558 / 35Normal (No Tools)-14.58
SWE-bench Multilingual78.304 / 25Thinking (With Tools)7215 / 25Extended (with tools)+6.30
SWE-bench Verified80.4013 / 114Thinking (With Tools)80.8410 / 114Extended (with tools)-0.44

AI Agent - Tool Usage

Even 2/2
BenchmarkQwen3.7 MaxClaude Opus 4.6Diff
Terminal Bench 2.069.705 / 48Thinking (With Tools)65.4011 / 48Extended (with tools)+4.30
MCP-Atlas76.4016 / 38Thinking (With Tools)76.8013 / 38Deep Thinking (With Tools)-0.40

General Knowledge

Even 2/2
BenchmarkQwen3.7 MaxClaude Opus 4.6Diff
LiveBench74.2921 / 115Deep Thinking (No Tools)76.338 / 115Thinking High (No Tools)-2.04
HLE53.5018 / 181Thinking (With Tools)5320 / 181Extended (with tools, internet)+0.50

Commonsense Reasoning

Qwen3.7 Max 1/1
BenchmarkQwen3.7 MaxClaude Opus 4.6Diff
SimpleBench70.407 / 67Normal (No Tools)67.609 / 67Normal (No Tools)+2.80

General Evaluation

Qwen3.7 Max 1/1
BenchmarkQwen3.7 MaxClaude Opus 4.6Diff
GPQA Diamond92.4022 / 226最高(无工具)91.3128 / 226Extended (no tools)+1.09

Instruction Following

Claude Opus 4.6 1/1
BenchmarkQwen3.7 MaxClaude Opus 4.6Diff
IF Bench79.104 / 33最高(无工具)941 / 33Extended (no tools)-14.90

Specs

FieldQwen3.7 MaxClaude Opus 4.6
Publisher阿里巴巴Anthropic
Release date2026-05-202026-02-05
Model typeReasoning modelReasoning model
ArchitectureDenseDense
ParametersNot availableNot available
Context length1M1000K
Max output64K64K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemQwen3.7 MaxClaude Opus 4.6
Text input¥12 / 1M tokens$0.5 / 1M tokens
Text output¥36 / 1M tokens$25 / 1M tokens
Cache readNot public$0.5 / 1M tokens
Cache writeNot public$10 / 1M tokens

Summary

  • Qwen3.7 Maxleads in:Commonsense Reasoning (1/1), General Evaluation (1/1)
  • Claude Opus 4.6leads in:Instruction Following (1/1)
  • Tied in:Coding and Software Engineer, AI Agent - Tool Usage, General Knowledge

On average across the 11 shared benchmarks, Claude Opus 4.6 scores 0.16 higher.

Largest single-benchmark gap: LiveCodeBench — Qwen3.7 Max 91.60 vs Claude Opus 4.6 76 (+15.60).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.