DataLearner logo

Qwen3.8-MaxvsQwen3.6-Max-Preview

Across 3 shared benchmarks, Qwen3.8-Max leads overall: Qwen3.8-Max wins 3, Qwen3.6-Max-Preview wins 0, with 0 ties and an average score difference of +6.20.

阿里巴巴
Qwen3.8-Max

阿里巴巴 · 2026-08-03 · Reasoning model

阿里巴巴
Qwen3.6-Max-Preview

阿里巴巴 · 2026-04-18 · Chat model

Qwen3.8-Max3 wins(100%)(0%)0 winsQwen3.6-Max-Preview

Benchmark scores

Grouped by capability, sorted by largest gap within each. 3 shared benchmarks.

Coding and Software Engineer

Qwen3.8-Max 1/1
BenchmarkQwen3.8-MaxQwen3.6-Max-PreviewDiff
SWE-Bench Pro - Public67.705 / 57极高强度思考(工具)57.3020 / 57Deep Thinking (With Tools)+10.40

General Evaluation

Qwen3.8-Max 1/1
BenchmarkQwen3.8-MaxQwen3.6-Max-PreviewDiff
GPQA Diamond92.6021 / 226极高强度思考(无工具)90.4036 / 226最高(无工具)+2.20

General Knowledge

Qwen3.8-Max 1/1
BenchmarkQwen3.8-MaxQwen3.6-Max-PreviewDiff
HLE56.2013 / 181极高强度思考(工具)50.2029 / 181Thinking (With Tools)+6

Specs

FieldQwen3.8-MaxQwen3.6-Max-Preview
Publisher阿里巴巴阿里巴巴
Release date2026-08-032026-04-18
Model typeReasoning modelChat model
ArchitectureMoEDense
Parameters2.4TNot available
Context length1M262K
Max output128K64K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemQwen3.8-MaxQwen3.6-Max-Preview
Text input¥12 / 1M tokens$1.3 / 1M tokens
Text output¥36 / 1M tokens$7.8 / 1M tokens
Cache read¥1.5 / 1M tokensNot public

Summary

  • Qwen3.8-Maxleads in:Coding and Software Engineer (1/1), General Evaluation (1/1), General Knowledge (1/1)

On average across the 3 shared benchmarks, Qwen3.8-Max scores 6.20 higher.

Largest single-benchmark gap: SWE-Bench Pro - Public — Qwen3.8-Max 67.70 vs Qwen3.6-Max-Preview 57.30 (+10.40).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.