DataLearner logo

Qwen3.6-27BvsGPT-5.4 mini

Across 7 shared benchmarks, Qwen3.6-27B leads overall: Qwen3.6-27B wins 6, GPT-5.4 mini wins 1, with 0 ties and an average score difference of +61.55.

阿里巴巴
Qwen3.6-27B

阿里巴巴 · 2026-04-22 · Reasoning model

OpenAI
GPT-5.4 mini

OpenAI · 2026-03-17 · Reasoning model

Qwen3.6-27B6 wins(86%)(14%)1 winGPT-5.4 mini

Benchmark scores

Grouped by capability, sorted by largest gap within each. 7 shared benchmarks.

Claw-style Agent Evaluation

GPT-5.4 mini 1/1
BenchmarkQwen3.6-27BGPT-5.4 miniDiff
Claw Bench72.4027 / 29Thinking (With Tools)75.3025 / 29Thinking (With Tools)-2.90

General Evaluation

Qwen3.6-27B 1/1
BenchmarkQwen3.6-27BGPT-5.4 miniDiff
GPQA Diamond84.85159 / 462Normal (No Tools)64.14221 / 274Normal (No Tools)+20.71

General Knowledge

Qwen3.6-27B 1/1
BenchmarkQwen3.6-27BGPT-5.4 miniDiff
LiveBench64.0354 / 117Normal (No Tools)36.95114 / 117Normal (No Tools)+27.08

Long Context

Qwen3.6-27B 1/1
BenchmarkQwen3.6-27BGPT-5.4 miniDiff
AA-LCR66.70118 / 170Normal (No Tools)37158 / 170Normal (No Tools)+29.70

Math and Reasoning

Qwen3.6-27B 1/1
BenchmarkQwen3.6-27BGPT-5.4 miniDiff
FrontierMath v234.0441 / 58Normal (No Tools)17.1954 / 58Normal (No Tools)+16.84

Productivity Knowledge

Qwen3.6-27B 1/1
BenchmarkQwen3.6-27BGPT-5.4 miniDiff
GDPval-AA v21,04387 / 105Normal (With Tools)734101 / 105Normal (With Tools)+309

Text Embedding

Qwen3.6-27B 1/1
BenchmarkQwen3.6-27BGPT-5.4 miniDiff
Context Arena53.7882 / 126Normal (No Tools)23.34119 / 126Normal (No Tools)+30.44

Specs

FieldQwen3.6-27BGPT-5.4 mini
Publisher阿里巴巴OpenAI
Release date2026-04-222026-03-17
Model typeReasoning modelReasoning model
ArchitectureDenseDense
Parameters27BNot available
Context length128K400K
Max output16K128K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemQwen3.6-27BGPT-5.4 mini
Text inputNot public$0.75 / 1M tokens
Text outputNot public$4.5 / 1M tokens
Cache readNot public$0.075 / 1M tokens

One or both models have incomplete public pricing.

Summary

  • Qwen3.6-27Bleads in:General Evaluation (1/1), General Knowledge (1/1), Long Context (1/1), Math and Reasoning (1/1), Productivity Knowledge (1/1), Text Embedding (1/1)
  • GPT-5.4 minileads in:Claw-style Agent Evaluation (1/1)

On average across the 7 shared benchmarks, Qwen3.6-27B scores 61.55 higher.

Largest single-benchmark gap: GDPval-AA v2 — Qwen3.6-27B 1,043 vs GPT-5.4 mini 734 (+309).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.