DataLearner logo

Qwen3.6-Max-PreviewvsOpus 4.7

Across 6 shared benchmarks, Opus 4.7 leads overall: Qwen3.6-Max-Preview wins 1, Opus 4.7 wins 5, with 0 ties and an average score difference of -4.67.

阿里巴巴
Qwen3.6-Max-Preview

阿里巴巴 · 2026-04-18 · Chat model

Anthropic
Opus 4.7

Anthropic · 2026-04-16 · Reasoning model

Qwen3.6-Max-Preview1 win(17%)(83%)5 winsOpus 4.7

Benchmark scores

Grouped by capability, sorted by largest gap within each. 6 shared benchmarks.

Coding and Software Engineer

Opus 4.7 2/2
BenchmarkQwen3.6-Max-PreviewOpus 4.7Diff
SWE-bench Verified78.8021 / 114Thinking (With Tools)87.606 / 114Extended (with tools)-8.80
SWE-Bench Pro - Public57.3020 / 57Deep Thinking (With Tools)64.308 / 57Extended (with tools)-7

AI Agent - Tool Usage

Opus 4.7 1/1
BenchmarkQwen3.6-Max-PreviewOpus 4.7Diff
Terminal Bench 2.065.4011 / 48Deep Thinking (With Tools)69.406 / 48Extended (with tools)-4

Commonsense Reasoning

Qwen3.6-Max-Preview 1/1
BenchmarkQwen3.6-Max-PreviewOpus 4.7Diff
SimpleBench6311 / 67Normal (No Tools)62.9012 / 67Normal (No Tools)+0.10

General Evaluation

Opus 4.7 1/1
BenchmarkQwen3.6-Max-PreviewOpus 4.7Diff
GPQA Diamond90.4036 / 226最高(无工具)94.205 / 226Extended (no tools)-3.80

General Knowledge

Opus 4.7 1/1
BenchmarkQwen3.6-Max-PreviewOpus 4.7Diff
HLE50.2029 / 181Thinking (With Tools)54.7015 / 181Extended (with tools)-4.50

Specs

FieldQwen3.6-Max-PreviewOpus 4.7
Publisher阿里巴巴Anthropic
Release date2026-04-182026-04-16
Model typeChat modelReasoning model
ArchitectureDenseDense
ParametersNot availableNot available
Context length262K1000K
Max output64K128K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemQwen3.6-Max-PreviewOpus 4.7
Text input$1.3 / 1M tokens$5 / 1M tokens
Text output$7.8 / 1M tokens$25 / 1M tokens
Cache readNot public$0.5 / 1M tokens
Cache writeNot public$6.25 / 1M tokens

Summary

  • Qwen3.6-Max-Previewleads in:Commonsense Reasoning (1/1)
  • Opus 4.7leads in:Coding and Software Engineer (2/2), AI Agent - Tool Usage (1/1), General Evaluation (1/1), General Knowledge (1/1)

On average across the 6 shared benchmarks, Opus 4.7 scores 4.67 higher.

Largest single-benchmark gap: SWE-bench Verified — Qwen3.6-Max-Preview 78.80 vs Opus 4.7 87.60 (-8.80).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.