DataLearner logo

Qwen3.6-27BvsHaiku 4.5

Across 4 shared benchmarks, Qwen3.6-27B leads overall: Qwen3.6-27B wins 3, Haiku 4.5 wins 1, with 0 ties and an average score difference of +15.54.

阿里巴巴
Qwen3.6-27B

阿里巴巴 · 2026-04-22 · Reasoning model

Anthropic
Haiku 4.5

Anthropic · 2025-10-15 · Multimodal model

Qwen3.6-27B3 wins(75%)(25%)1 winHaiku 4.5

Benchmark scores

Grouped by capability, sorted by largest gap within each. 4 shared benchmarks.

Claw-style Agent Evaluation

Haiku 4.5 1/1
BenchmarkQwen3.6-27BHaiku 4.5Diff
Claw Bench72.4027 / 29Thinking (With Tools)89.4011 / 29Thinking (With Tools)-17

General Evaluation

Qwen3.6-27B 1/1
BenchmarkQwen3.6-27BHaiku 4.5Diff
GPQA Diamond84.85103 / 274Normal (No Tools)60.50227 / 274Normal (No Tools)+24.35

General Knowledge

Qwen3.6-27B 1/1
BenchmarkQwen3.6-27BHaiku 4.5Diff
LiveBench64.0354 / 117Normal (No Tools)45.33105 / 117Normal (No Tools)+18.70

Text Embedding

Qwen3.6-27B 1/1
BenchmarkQwen3.6-27BHaiku 4.5Diff
Context Arena53.7882 / 126Normal (No Tools)17.68124 / 126Normal (No Tools)+36.10

Specs

FieldQwen3.6-27BHaiku 4.5
Publisher阿里巴巴Anthropic
Release date2026-04-222025-10-15
Model typeReasoning modelMultimodal model
ArchitectureDenseDense
Parameters27BNot available
Context length128K200K
Max output16K64K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemQwen3.6-27BHaiku 4.5
Text inputNot public$1 / 1M tokens
Text outputNot public$5 / 1M tokens
Cache readNot public$0.1 / 1M tokens
Cache writeNot public$1.25 / 1M tokens

One or both models have incomplete public pricing.

Summary

  • Qwen3.6-27Bleads in:General Evaluation (1/1), General Knowledge (1/1), Text Embedding (1/1)
  • Haiku 4.5leads in:Claw-style Agent Evaluation (1/1)

On average across the 4 shared benchmarks, Qwen3.6-27B scores 15.54 higher.

Largest single-benchmark gap: Context Arena — Qwen3.6-27B 53.78 vs Haiku 4.5 17.68 (+36.10).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.