DataLearner logo

Qwen3.7 MaxvsQwen3-Max-Thinking

Across 7 shared benchmarks, Qwen3.7 Max leads overall: Qwen3.7 Max wins 7, Qwen3-Max-Thinking wins 0, with 0 ties and an average score difference of +5.39.

阿里巴巴
Qwen3.7 Max

阿里巴巴 · 2026-05-20 · Reasoning model

阿里巴巴
Qwen3-Max-Thinking

阿里巴巴 · 2026-01-26 · Reasoning model

Qwen3.7 Max7 wins(100%)(0%)0 winsQwen3-Max-Thinking

Benchmark scores

Grouped by capability, sorted by largest gap within each. 7 shared benchmarks.

Coding and Software Engineer

Qwen3.7 Max 2/2
BenchmarkQwen3.7 MaxQwen3-Max-ThinkingDiff
LiveCodeBench91.604 / 126最高(无工具)85.9016 / 126+5.70
SWE-bench Verified80.4013 / 114Thinking (With Tools)75.3038 / 114+5.10

General Knowledge

Qwen3.7 Max 2/2
BenchmarkQwen3.7 MaxQwen3-Max-ThinkingDiff
MMLU Pro89.604 / 133最高(无工具)85.7023 / 133+3.90
HLE53.5018 / 181Thinking (With Tools)49.8032 / 181+3.70

General Evaluation

Qwen3.7 Max 1/1
BenchmarkQwen3.7 MaxQwen3-Max-ThinkingDiff
GPQA Diamond92.4022 / 226最高(无工具)87.4066 / 226+5

Instruction Following

Qwen3.7 Max 1/1
BenchmarkQwen3.7 MaxQwen3-Max-ThinkingDiff
IF Bench79.104 / 33最高(无工具)70.9014 / 33+8.20

Math and Reasoning

Qwen3.7 Max 1/1
BenchmarkQwen3.7 MaxQwen3-Max-ThinkingDiff
IMO-AnswerBench903 / 23最高(无工具)83.9013 / 23+6.10

Specs

FieldQwen3.7 MaxQwen3-Max-Thinking
Publisher阿里巴巴阿里巴巴
Release date2026-05-202026-01-26
Model typeReasoning modelReasoning model
ArchitectureDenseMoE
ParametersNot available1T
Context length1M1000K
Max output64K32K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemQwen3.7 MaxQwen3-Max-Thinking
Text input¥12 / 1M tokens$1.2 / 1M tokens
Text output¥36 / 1M tokens$6 / 1M tokens

Summary

  • Qwen3.7 Maxleads in:Coding and Software Engineer (2/2), General Knowledge (2/2), General Evaluation (1/1), Instruction Following (1/1), Math and Reasoning (1/1)

On average across the 7 shared benchmarks, Qwen3.7 Max scores 5.39 higher.

Largest single-benchmark gap: IF Bench — Qwen3.7 Max 79.10 vs Qwen3-Max-Thinking 70.90 (+8.20).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.