DataLearner logo

Qwen3.8-27B Benchmark Details

Qwen3.8-27B currently shows benchmark results led by LiveCodeBench (6 / 124, score 90.30), OSWorld-Verified (3 / 26, score 84.30), CharXiv RQ (2 / 15, score 90.20). This page also compares it with 4 competitor models and 4 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

Qwen3.8-27B

Benchmark Results

Thinking
Tool usage

General Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
HLE
Thinking Mode
30.80
87 / 179

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Mode
89.20
47 / 226

Coding and Software Engineer

4 evaluations
Benchmark / mode
Score
Rank/total
LiveCodeBench
Thinking Mode
90.30
6 / 124
SWE-Bench Pro - Public
Thinking ModeTools
61.70
10 / 57
NL2Repo-Bench
Thinking ModeTools
42.30
4 / 4
DeepSWE
Thinking ModeTools
42.20
20 / 25

AI Agent - Tool Usage

2 evaluations
Benchmark / mode
Score
Rank/total
OSWorld-Verified
Thinking ModeTools
84.30
3 / 26
Terminal-Bench 2.1
Thinking ModeTools
73
27 / 41

Agent Level Benchmark

1 evaluations
Benchmark / mode
Score
Rank/total
Agents' Last Exam
Thinking ModeTools
20.40
9 / 9

Multimodal Understanding

7 evaluations
Benchmark / mode
Score
Rank/total
MathVision
Thinking Mode
90
6 / 6
MathVision
Thinking ModeTools
94.60
2 / 6
OmniDocBench
Thinking Mode
91.10
1 / 3
CharXiv RQ
Thinking Mode
83.70
12 / 15
CharXiv RQ
Thinking ModeTools
90.20
2 / 15
BabyVision
Thinking Mode
65.70
4 / 4
BabyVision
Thinking ModeTools
85.60
2 / 4

Competitor Comparison

Benchmark scores for Qwen3.8-27B compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

10 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkQwen3.8-27BCurrentClaude Sonnet 5GPT-5.6 TerraGemini 3.7 FlashDeepSeek-V4-Flash
HLE
综合评估
30.80Thinking Enabled
57.40Thinking Level · Extra High | Tools
--
--
45.10Thinking Level · Extra High | Tools
GPQA Diamond
科学与综合推理
89.20Thinking Enabled
90.53Thinking Level · Extra High
93.31Thinking Level · High
--
88.10Thinking Level · High
DeepSWE
编程与软件工程
42.20Thinking Enabled | Tools
54.00Deep Thinking Mode | Tools
69.60Thinking Level · Extra High | Tools
65.30Thinking Level · High | Tools
54.40Thinking Level · High | Tools
LiveCodeBench
编程与软件工程
90.30Thinking Enabled
--
--
--
91.60Thinking Level · High
NL2Repo-Bench
编程与软件工程
42.30Thinking Enabled | Tools
--
--
--
54.20Thinking Level · High | Tools
SWE-Bench Pro - Public
编程与软件工程
61.70Thinking Enabled | Tools
--
--
--
52.60Thinking Level · Extra High | Tools
OSWorld-Verified
AI Agent - 工具使用
84.30Thinking Enabled | Tools
81.20Thinking Level · Extra High | Tools
--
--
--
Terminal-Bench 2.1
AI Agent - 工具使用
73.00Thinking Enabled | Tools
80.40Thinking Level · Extra High | Tools
87.40Thinking Level · High
85.80Thinking Enabled | Tools
82.70Thinking Level · High | Tools
Agents' Last Exam
Agent能力评测
20.40Thinking Enabled | Tools
--
50.40Thinking Level · Extra High | Tools
26.30Thinking Level · Medium | Tools
25.20Thinking Level · High | Tools
CharXiv RQ
多模态理解
90.20Thinking Enabled | Tools
--
--
88.70Thinking Level · Medium | Tools
--

Standard API Pricing: Qwen3.8-27B vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Claude Sonnet 5
Anthropic$2 / 1M tokens$10 / 1M tokens
GPT-5.6 Terra
OpenAI$2.5 / 1M tokens$15 / 1M tokens
Gemini 3.7 Flash
DeepMind$0.75 / 1M tokens$3.75 / 1M tokens
DeepSeek-V4-Flash
DeepSeek-AI$0.14 / 1M tokens$0.28 / 1M tokens

Version History

How each version of the Qwen3.8-27B series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

5 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkQwen3.8-27BCurrentQwen3.6-27BQwen3.5-27BQwen3-32BQwen2.5-32B
HLE
综合评估
30.80Thinking Enabled
24.00Thinking Enabled
48.50Thinking Enabled | Tools
--
--
GPQA Diamond
科学与综合推理
89.20Thinking Enabled
87.80Thinking Enabled
85.50Thinking Enabled
68.40Thinking Enabled
--
LiveCodeBench
编程与软件工程
90.30Thinking Enabled
83.90Thinking Enabled
80.70Thinking Enabled | Tools
65.70Thinking Enabled
51.20Standard Mode
SWE-Bench Pro - Public
编程与软件工程
61.70Thinking Enabled | Tools
53.50Thinking Enabled | Tools
--
--
--
OSWorld-Verified
AI Agent - 工具使用
84.30Thinking Enabled | Tools
--
56.20Thinking Enabled | Tools
--
--

Single-Benchmark Version Trend

Viewing: HLE · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Qwen3.8-27B Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

Qwen3-32B
Supplier: 阿里巴巴
Standard input: ¥0.0012 / 1K tokens
Standard output: ¥0.0048 / 1K tokens
Qwen2.5-32B
Supplier: 阿里巴巴
Standard input: ¥0.002 / 1K tokens
Standard output: ¥0.006 / 1K tokens
ModelSupplierStandard inputStandard outputBase price applies to
Qwen3-32B
阿里巴巴¥0.0012 / 1K tokens¥0.0048 / 1K tokens
Qwen2.5-32B
阿里巴巴¥0.002 / 1K tokens¥0.006 / 1K tokens