DataLearner logo

Qwen3.8-27BvsDeepSeek-V4-Flash

Qwen3.8-27B and DeepSeek-V4-Flash are tied across 8 shared benchmarks: Qwen3.8-27B leads on 4, DeepSeek-V4-Flash leads on 4, with 0 ties and an average score difference of +6.23.

阿里巴巴
Qwen3.8-27B

阿里巴巴 · 2026-08-14 · Reasoning model

DeepSeek-AI
DeepSeek-V4-Flash

DeepSeek-AI · 2026-04-24 · Reasoning model

Qwen3.8-27B4 wins(50%)(50%)4 winsDeepSeek-V4-Flash

Benchmark scores

Grouped by capability, sorted by largest gap within each. 8 shared benchmarks.

Coding and Software Engineer

Even 4/4
BenchmarkQwen3.8-27BDeepSeek-V4-FlashDiff
LiveCodeBench90.306 / 124Thinking (No Tools)55.2084 / 124Normal (No Tools)+35.10
SWE-Bench Pro - Public61.7010 / 57Thinking (With Tools)49.1047 / 57Normal (With Tools)+12.60
DeepSWE42.2020 / 25Thinking (With Tools)54.4013 / 25Max (With Tools)-12.20
NL2Repo-Bench42.304 / 4Thinking (With Tools)54.203 / 4Max (With Tools)-11.90

Agent Level Benchmark

DeepSeek-V4-Flash 1/1
BenchmarkQwen3.8-27BDeepSeek-V4-FlashDiff
Agents' Last Exam20.409 / 9Thinking (With Tools)25.208 / 9Max (With Tools)-4.80

AI Agent - Tool Usage

DeepSeek-V4-Flash 1/1
BenchmarkQwen3.8-27BDeepSeek-V4-FlashDiff
Terminal-Bench 2.17327 / 41Thinking (With Tools)82.7014 / 41Max (With Tools)-9.70

General Evaluation

Qwen3.8-27B 1/1
BenchmarkQwen3.8-27BDeepSeek-V4-FlashDiff
GPQA Diamond89.2047 / 226Thinking (No Tools)71.20151 / 226Normal (No Tools)+18

General Knowledge

Qwen3.8-27B 1/1
BenchmarkQwen3.8-27BDeepSeek-V4-FlashDiff
HLE30.8087 / 179Thinking (No Tools)8.10162 / 179Normal (No Tools)+22.70

Specs

FieldQwen3.8-27BDeepSeek-V4-Flash
Publisher阿里巴巴DeepSeek-AI
Release date2026-08-142026-04-24
Model typeReasoning modelReasoning model
ArchitectureDenseMoE
Parameters27B284B
Context length256K1M
Max output128K384K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemQwen3.8-27BDeepSeek-V4-Flash
Text inputNot public$0.14 / 1M tokens
Text outputNot public$0.28 / 1M tokens
Cache readNot public$0.0028 / 1M tokens

One or both models have incomplete public pricing.

Summary

  • Qwen3.8-27Bleads in:General Evaluation (1/1), General Knowledge (1/1)
  • DeepSeek-V4-Flashleads in:Agent Level Benchmark (1/1), AI Agent - Tool Usage (1/1)
  • Tied in:Coding and Software Engineer

On average across the 8 shared benchmarks, Qwen3.8-27B scores 6.23 higher.

Largest single-benchmark gap: LiveCodeBench — Qwen3.8-27B 90.30 vs DeepSeek-V4-Flash 55.20 (+35.10).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.