DataLearner logo

Qwen3.8-MaxvsClaude Opus 5

Across 7 shared benchmarks, Claude Opus 5 leads overall: Qwen3.8-Max wins 1, Claude Opus 5 wins 6, with 0 ties and an average score difference of -9.98.

阿里巴巴
Qwen3.8-Max

阿里巴巴 · 2026-08-03 · Reasoning model

Anthropic
Claude Opus 5

Anthropic · 2026-07-24 · Reasoning model

Qwen3.8-Max1 win(14%)(86%)6 winsClaude Opus 5

Benchmark scores

Grouped by capability, sorted by largest gap within each. 7 shared benchmarks.

Coding and Software Engineer

Claude Opus 5 2/2
BenchmarkQwen3.8-MaxClaude Opus 5Diff
DeepSWE56.6014 / 27极高强度思考(工具)68.804 / 27Max (With Tools)-12.20
SWE-Bench Pro - Public67.705 / 57极高强度思考(工具)79.202 / 57Max (With Tools)-11.50

Math and Reasoning

Claude Opus 5 2/2
BenchmarkQwen3.8-MaxClaude Opus 5Diff
FrontierMath Tier 4 v246.3411 / 34极高强度思考(无工具)73.174 / 34最高(无工具)-26.83
FrontierMath v274.7411 / 34极高强度思考(无工具)85.615 / 34最高(无工具)-10.88

AI Agent - Tool Usage

Qwen3.8-Max 1/1
BenchmarkQwen3.8-MaxClaude Opus 5Diff
AutomationBench27.305 / 8极高强度思考(工具)266 / 8Max (With Tools)+1.30

General Evaluation

Claude Opus 5 1/1
BenchmarkQwen3.8-MaxClaude Opus 5Diff
GPQA Diamond92.6021 / 226极高强度思考(无工具)93.889 / 226最高(无工具)-1.28

General Knowledge

Claude Opus 5 1/1
BenchmarkQwen3.8-MaxClaude Opus 5Diff
HLE56.2013 / 181极高强度思考(工具)64.701 / 181Max (With Tools)-8.50

Specs

FieldQwen3.8-MaxClaude Opus 5
Publisher阿里巴巴Anthropic
Release date2026-08-032026-07-24
Model typeReasoning modelReasoning model
ArchitectureMoEDense
Parameters2.4TNot available
Context length1M1M
Max output128K128K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemQwen3.8-MaxClaude Opus 5
Text input¥12 / 1M tokens$5 / 1M tokens
Text output¥36 / 1M tokens$25 / 1M tokens
Cache read¥1.5 / 1M tokens$0.5 / 1M tokens
Cache writeNot public$6.25 / 1M tokens

Summary

  • Qwen3.8-Maxleads in:AI Agent - Tool Usage (1/1)
  • Claude Opus 5leads in:Coding and Software Engineer (2/2), Math and Reasoning (2/2), General Evaluation (1/1), General Knowledge (1/1)

On average across the 7 shared benchmarks, Claude Opus 5 scores 9.98 higher.

Largest single-benchmark gap: FrontierMath Tier 4 v2 — Qwen3.8-Max 46.34 vs Claude Opus 5 73.17 (-26.83).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.