DataLearner logo

Qwen3.8-27BvsClaude Sonnet 5

Across 5 shared benchmarks, Claude Sonnet 5 leads overall: Qwen3.8-27B wins 1, Claude Sonnet 5 wins 4, with 0 ties and an average score difference of -8.81.

阿里巴巴
Qwen3.8-27B

阿里巴巴 · 2026-08-14 · Reasoning model

Anthropic
Claude Sonnet 5

Anthropic · 2026-06-30 · Multimodal model

Qwen3.8-27B1 win(20%)(80%)4 winsClaude Sonnet 5

Benchmark scores

Grouped by capability, sorted by largest gap within each. 5 shared benchmarks.

AI Agent - Tool Usage

Even 2/2
BenchmarkQwen3.8-27BClaude Sonnet 5Diff
Terminal-Bench 2.17327 / 41Thinking (With Tools)80.4016 / 41极高强度思考(工具)-7.40
OSWorld-Verified84.303 / 26Thinking (With Tools)81.206 / 26极高强度思考(工具)+3.10

Coding and Software Engineer

Claude Sonnet 5 1/1
BenchmarkQwen3.8-27BClaude Sonnet 5Diff
DeepSWE42.2020 / 25Thinking (With Tools)5414 / 25Deep Thinking (With Tools)-11.80

General Evaluation

Claude Sonnet 5 1/1
BenchmarkQwen3.8-27BClaude Sonnet 5Diff
GPQA Diamond89.2047 / 226Thinking (No Tools)90.5333 / 226极高强度思考(无工具)-1.33

General Knowledge

Claude Sonnet 5 1/1
BenchmarkQwen3.8-27BClaude Sonnet 5Diff
HLE30.8087 / 179Thinking (No Tools)57.409 / 179极高强度思考(工具)-26.60

Specs

FieldQwen3.8-27BClaude Sonnet 5
Publisher阿里巴巴Anthropic
Release date2026-08-142026-06-30
Model typeReasoning modelMultimodal model
ArchitectureDenseDense
Parameters27BNot available
Context length256K1M
Max output128K128K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemQwen3.8-27BClaude Sonnet 5
Text inputNot public$2 / 1M tokens
Text outputNot public$10 / 1M tokens
Cache readNot public$0.3 / 1M tokens
Cache writeNot public$3.75 / 1M tokens

One or both models have incomplete public pricing.

Summary

  • Claude Sonnet 5leads in:Coding and Software Engineer (1/1), General Evaluation (1/1), General Knowledge (1/1)
  • Tied in:AI Agent - Tool Usage

On average across the 5 shared benchmarks, Claude Sonnet 5 scores 8.81 higher.

Largest single-benchmark gap: HLE — Qwen3.8-27B 30.80 vs Claude Sonnet 5 57.40 (-26.60).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.