DataLearner logo

DeepSeek-V4-FlashvsQwen3.6-27B

Across 10 shared benchmarks, Qwen3.6-27B leads overall: DeepSeek-V4-Flash wins 1, Qwen3.6-27B wins 9, with 0 ties and an average score difference of -12.13.

DeepSeek-AI
DeepSeek-V4-Flash

DeepSeek-AI · 2026-04-24 · Reasoning model

阿里巴巴
Qwen3.6-27B

阿里巴巴 · 2026-04-22 · Reasoning model

DeepSeek-V4-Flash1 win(10%)(90%)9 winsQwen3.6-27B

Benchmark scores

Grouped by capability, sorted by largest gap within each. 10 shared benchmarks.

Coding and Software Engineer

Qwen3.6-27B 4/4
BenchmarkDeepSeek-V4-FlashQwen3.6-27BDiff
LiveCodeBench55.2086 / 126Normal (No Tools)83.9021 / 126Thinking (No Tools)-28.70
SWE-Bench Pro - Public49.1047 / 57Normal (With Tools)53.5036 / 57Thinking (With Tools)-4.40
SWE-bench Verified73.7045 / 114Normal (With Tools)77.2028 / 114Thinking (With Tools)-3.50
SWE-bench Multilingual69.7020 / 25Normal (With Tools)71.3017 / 25Thinking (With Tools)-1.60

General Knowledge

Qwen3.6-27B 2/3
BenchmarkDeepSeek-V4-FlashQwen3.6-27BDiff
HLE8.10164 / 181Normal (No Tools)24115 / 181Thinking (No Tools)-15.90
MMLU Pro8346 / 133Normal (No Tools)86.2018 / 133Thinking (No Tools)-3.20
LiveBench67.2549 / 115Normal (No Tools)65.5652 / 115Normal (No Tools)+1.69

AI Agent - Tool Usage

Qwen3.6-27B 1/1
BenchmarkDeepSeek-V4-FlashQwen3.6-27BDiff
Terminal Bench 2.049.1036 / 48Normal (With Tools)59.3020 / 48Thinking (With Tools)-10.20

General Evaluation

Qwen3.6-27B 1/1
BenchmarkDeepSeek-V4-FlashQwen3.6-27BDiff
GPQA Diamond71.20151 / 226Normal (No Tools)87.8064 / 226Thinking (No Tools)-16.60

Math and Reasoning

Qwen3.6-27B 1/1
BenchmarkDeepSeek-V4-FlashQwen3.6-27BDiff
IMO-AnswerBench41.9022 / 23Normal (No Tools)80.8020 / 23Thinking (No Tools)-38.90

Specs

FieldDeepSeek-V4-FlashQwen3.6-27B
PublisherDeepSeek-AI阿里巴巴
Release date2026-04-242026-04-22
Model typeReasoning modelReasoning model
ArchitectureMoEDense
Parameters284B27B
Context length1M128K
Max output384K16K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemDeepSeek-V4-FlashQwen3.6-27B
Text input$0.14 / 1M tokensNot public
Text output$0.28 / 1M tokensNot public
Cache read$0.0028 / 1M tokensNot public

One or both models have incomplete public pricing.

Summary

  • Qwen3.6-27Bleads in:Coding and Software Engineer (4/4), General Knowledge (2/3), AI Agent - Tool Usage (1/1), General Evaluation (1/1), Math and Reasoning (1/1)

On average across the 10 shared benchmarks, Qwen3.6-27B scores 12.13 higher.

Largest single-benchmark gap: IMO-AnswerBench — DeepSeek-V4-Flash 41.90 vs Qwen3.6-27B 80.80 (-38.90).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.