DataLearner logo

Qwen3.8-MaxvsGPT-5.6 Sol

Across 7 shared benchmarks, GPT-5.6 Sol leads overall: Qwen3.8-Max wins 1, GPT-5.6 Sol wins 6, with 0 ties and an average score difference of -13.25.

阿里巴巴
Qwen3.8-Max

阿里巴巴 · 2026-08-03 · Reasoning model

OpenAI
GPT-5.6 Sol

OpenAI · 2026-06-26 · Reasoning model

Qwen3.8-Max1 win(14%)(86%)6 winsGPT-5.6 Sol

Benchmark scores

Grouped by capability, sorted by largest gap within each. 7 shared benchmarks.

Coding and Software Engineer

Even 2/2
BenchmarkQwen3.8-MaxGPT-5.6 SolDiff
DeepSWE56.6014 / 27极高强度思考(工具)72.701 / 27极高强度思考(工具)-16.10
SWE-Bench Pro - Public67.705 / 57极高强度思考(工具)64.607 / 57极高强度思考(工具)+3.10

Math and Reasoning

GPT-5.6 Sol 2/2
BenchmarkQwen3.8-MaxGPT-5.6 SolDiff
FrontierMath Tier 4 v246.3411 / 34极高强度思考(无工具)82.932 / 34最高(无工具)-36.59
FrontierMath v274.7411 / 34极高强度思考(无工具)89.121 / 34最高(无工具)-14.39

Agent Level Benchmark

GPT-5.6 Sol 1/1
BenchmarkQwen3.8-MaxGPT-5.6 SolDiff
Agents' Last Exam277 / 11极高强度思考(工具)52.701 / 11极高强度思考(工具)-25.70

AI Agent - Tool Usage

GPT-5.6 Sol 1/1
BenchmarkQwen3.8-MaxGPT-5.6 SolDiff
Terminal-Bench 2.186.608 / 44极高强度思考(工具)88.801 / 44最高(无工具)-2.20

General Evaluation

GPT-5.6 Sol 1/1
BenchmarkQwen3.8-MaxGPT-5.6 SolDiff
GPQA Diamond92.6021 / 226极高强度思考(无工具)93.5014 / 226最高(无工具)-0.90

Specs

FieldQwen3.8-MaxGPT-5.6 Sol
Publisher阿里巴巴OpenAI
Release date2026-08-032026-06-26
Model typeReasoning modelReasoning model
ArchitectureMoEDense
Parameters2.4TNot available
Context length1M1.05M
Max output128K128K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemQwen3.8-MaxGPT-5.6 Sol
Text input¥12 / 1M tokens$5 / 1M tokens
Text output¥36 / 1M tokens$30 / 1M tokens
Cache read¥1.5 / 1M tokens$0.5 / 1M tokens
Cache writeNot public$6.25 / 1M tokens

Summary

  • GPT-5.6 Solleads in:Math and Reasoning (2/2), Agent Level Benchmark (1/1), AI Agent - Tool Usage (1/1), General Evaluation (1/1)
  • Tied in:Coding and Software Engineer

On average across the 7 shared benchmarks, GPT-5.6 Sol scores 13.25 higher.

Largest single-benchmark gap: FrontierMath Tier 4 v2 — Qwen3.8-Max 46.34 vs GPT-5.6 Sol 82.93 (-36.59).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.