DataLearner logo

Qwen3.8-MaxvsQwen3.7 Max

Across 3 shared benchmarks, Qwen3.8-Max leads overall: Qwen3.8-Max wins 3, Qwen3.7 Max wins 0, with 0 ties and an average score difference of +3.33.

阿里巴巴
Qwen3.8-Max

阿里巴巴 · 2026-08-03 · Reasoning model

阿里巴巴
Qwen3.7 Max

阿里巴巴 · 2026-05-20 · Reasoning model

Qwen3.8-Max3 wins(100%)(0%)0 winsQwen3.7 Max

Benchmark scores

Grouped by capability, sorted by largest gap within each. 3 shared benchmarks.

Coding and Software Engineer

Qwen3.8-Max 1/1
BenchmarkQwen3.8-MaxQwen3.7 MaxDiff
SWE-Bench Pro - Public67.705 / 57极高强度思考(工具)60.6012 / 57Thinking (With Tools)+7.10

General Evaluation

Qwen3.8-Max 1/1
BenchmarkQwen3.8-MaxQwen3.7 MaxDiff
GPQA Diamond92.6021 / 226极高强度思考(无工具)92.4022 / 226最高(无工具)+0.20

General Knowledge

Qwen3.8-Max 1/1
BenchmarkQwen3.8-MaxQwen3.7 MaxDiff
HLE56.2013 / 181极高强度思考(工具)53.5018 / 181Thinking (With Tools)+2.70

Specs

FieldQwen3.8-MaxQwen3.7 Max
Publisher阿里巴巴阿里巴巴
Release date2026-08-032026-05-20
Model typeReasoning modelReasoning model
ArchitectureMoEDense
Parameters2.4TNot available
Context length1M1M
Max output128K64K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemQwen3.8-MaxQwen3.7 Max
Text input¥12 / 1M tokens¥12 / 1M tokens
Text output¥36 / 1M tokens¥36 / 1M tokens
Cache read¥1.5 / 1M tokensNot public

Summary

  • Qwen3.8-Maxleads in:Coding and Software Engineer (1/1), General Evaluation (1/1), General Knowledge (1/1)

On average across the 3 shared benchmarks, Qwen3.8-Max scores 3.33 higher.

Largest single-benchmark gap: SWE-Bench Pro - Public — Qwen3.8-Max 67.70 vs Qwen3.7 Max 60.60 (+7.10).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.