DataLearner logo

Qwen3.8-27B Benchmark Details

Qwen3.8-27B currently shows benchmark results led by τ³-Banking (7 / 167, score 48), Context Arena (7 / 126, score 93.98), LiveCodeBench (7 / 125, score 90.30). This page also compares it with 4 competitor models and 4 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

Qwen3.8-27B

Benchmark Results

Thinking
Tool usage

Knowledge Exams

1 evaluations
Benchmark / mode
Score
Rank/total
HLE
Thinking Mode
30.80
112 / 235

Scientific Reasoning

3 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Mode
89.20
49 / 253
CritPt
Standard Mode
0.30
186 / 204
CritPt
Thinking Level · Extra High
5.40
96 / 204

Algorithmic Coding

2 evaluations
Benchmark / mode
Score
Rank/total
LiveCodeBench
Thinking Mode
90.30
7 / 125
IOI (Vals v2)
Thinking Level · Extra HighTools
39.06
24 / 26

Writing

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1668.40
32 / 106

Commonsense

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Thinking Mode
60.20
37 / 96

Repository Engineering

3 evaluations
Benchmark / mode
Score
Rank/total
SWE-Bench Pro - Public
Thinking ModeTools
61.70
13 / 62
NL2Repo-Bench
Thinking ModeTools
42.30
15 / 16
DeepSWE
Thinking ModeTools
42.20
75 / 91

Memory & Persistence

4 evaluations
Benchmark / mode
Score
Rank/total
Context Arena
Standard Mode
46.62
93 / 126
Context Arena
Thinking Level · Low
92.79
9 / 126
Context Arena
Thinking Level · Medium
93.98
7 / 126
Context Arena
Thinking Level · Extra High
90.17
14 / 126

Desktop Workflows

1 evaluations
Benchmark / mode
Score
Rank/total
OSWorld-Verified
Thinking ModeTools
84.30
3 / 28

Agentic Development

4 evaluations
Benchmark / mode
Score
Rank/total
Terminal-Bench 2.1
Thinking ModeTools
73
39 / 57
Terminal-Bench 4.0
Thinking Level · LowTools
2.50
74 / 97
Terminal-Bench 4.0
Thinking Level · MediumTools
5.10
66 / 97
Terminal-Bench 4.0
Thinking Level · Extra HighTools
5.60
65 / 97

Capability Indices

2 evaluations
Benchmark / mode
Score
Rank/total
ECI
unknown
149.38
50 / 167
Vals Index
Thinking Level · Extra HighTools
48.48
32 / 42

Tool Orchestration

1 evaluations
Benchmark / mode
Score
Rank/total
Agents' Last Exam
Thinking ModeTools
20.40
24 / 24

Visual Understanding

8 evaluations
Benchmark / mode
Score
Rank/total
MathVision
Thinking Mode
90
8 / 12
MathVision
Thinking ModeTools
94.60
3 / 12
BabyVision
Thinking Mode
65.70
8 / 9
BabyVision
Thinking ModeTools
85.60
6 / 9
MMMU-Pro
Standard Mode
69.90
133 / 229
MMMU-Pro
Thinking Level · Low
73.80
104 / 229
MMMU-Pro
Thinking Level · Medium
74.20
97 / 229
MMMU-Pro
Thinking Level · Extra High
76.30
75 / 229

Documents & Charts

6 evaluations
Benchmark / mode
Score
Rank/total
OmniDocBench
Thinking Mode
91.10
1 / 3
CharXiv RQ
Thinking Mode
83.70
16 / 19
CharXiv RQ
Thinking ModeTools
90.20
3 / 19
GDP.pdf
Thinking Level · Low
10.40
82 / 122
GDP.pdf
Thinking Level · Medium
12.80
70 / 122
GDP.pdf
Thinking Level · Extra High
16.40
56 / 122

Scientific Computing

4 evaluations
Benchmark / mode
Score
Rank/total
SciCode
Standard Mode
36.20
118 / 134
SciCode
Thinking Level · Low
40
106 / 134
SciCode
Thinking Level · Medium
39
110 / 134
SciCode
Thinking Level · Extra High
46.60
86 / 134

Service Workflows

5 evaluations
Benchmark / mode
Score
Rank/total
SAGE
Thinking Level · Extra HighTools
52.40
5 / 64
τ³-Banking
Standard ModeTools
20
94 / 167
τ³-Banking
Thinking Level · LowTools
32.20
54 / 167
τ³-Banking
Thinking Level · MediumTools
47.40
9 / 167
τ³-Banking
Thinking Level · Extra HighTools
48
7 / 167

Finance

2 evaluations
Benchmark / mode
Score
Rank/total
EMB (Excel Modeling)
Thinking Level · Extra HighTools
59.66
26 / 42
Finance Agent v2
Thinking Level · Extra HighTools
48.55
32 / 42

Legal

2 evaluations
Benchmark / mode
Score
Rank/total
Legal Research Bench
Thinking Level · Extra HighTools
36.06
27 / 42
Harvey's Legal Agent Benchmark
Thinking Level · Extra HighTools
11.25
8 / 43

Maintenance & Optimization

1 evaluations
Benchmark / mode
Score
Rank/total
Code Migration
Thinking Level · Extra HighTools
14.16
36 / 43

Code Generation & Editing

1 evaluations
Benchmark / mode
Score
Rank/total
Vibe Code Bench v1.1
Thinking Level · Extra HighTools
64.85
28 / 60

Mathematics

1 evaluations
Benchmark / mode
Score
Rank/total
ProofBench v1.1 (Lean 4)
Thinking Level · Extra HighTools
16
27 / 28

Clinical Workflows

2 evaluations
Benchmark / mode
Score
Rank/total
MedScribe
Thinking Level · Extra HighTools
83.85
28 / 66
MedCode
Thinking Level · Extra HighTools
28.70
61 / 64

Competitor Comparison

Benchmark scores for Qwen3.8-27B compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkQwen3.8-27BCurrentClaude Sonnet 5GPT-5.6 TerraGemini 3.7 FlashDeepSeek-V4-Flash
HLE
Accuracy
Knowledge Exams
30.80Thinking Enabled
57.40Thinking Level · Extra High | Tools
--
--
51.50Thinking Level · High | Tools
CritPt
Score
Scientific Reasoning
5.40Thinking Level · Extra High
16.90Thinking Level · High
30.00Thinking Level · High
14.30Thinking Level · High
7.10Thinking Level · High
GPQA Diamond
Accuracy
Scientific Reasoning
89.20Thinking Enabled
90.53Thinking Level · Extra High
93.31Thinking Level · High
94.82Thinking Level · High
71.20Standard Mode
IOI (Vals v2)
Subtask points (%)
Algorithmic Coding
39.06Thinking Level · Extra High | Tools
45.00Thinking Level · High | Tools
87.61Thinking Level · High | Tools
67.83Thinking Level · High | Tools
--
LiveCodeBench
Pass @K
Algorithmic Coding
90.30Thinking Enabled
--
--
--
91.60Thinking Level · High
Creative Writing
Elo、大模型评判两两对战
Writing
1668.40Standard Mode
1790.50Standard Mode
1850.30Standard Mode
1722.00Standard Mode
1555.70Standard Mode
SimpleBench
Score (AVG@5)
Commonsense
60.20Thinking Enabled
60.60Standard Mode
48.90Thinking Level · Extra High
--
61.10Standard Mode
DeepSWE
Pass@1 (DeepSWE v1.1)
Repository Engineering
42.20Thinking Enabled | Tools
54.00Deep Thinking Mode | Tools
69.62Thinking Level · High | Tools
65.49Thinking Level · Medium | Tools
53.32Thinking Level · High | Tools
NL2Repo-Bench
Average test pass rate
Repository Engineering
42.30Thinking Enabled | Tools
--
--
--
54.20Thinking Level · High | Tools
SWE-Bench Pro - Public
Accuracy
Repository Engineering
61.70Thinking Enabled | Tools
--
--
--
52.60Thinking Level · Extra High | Tools
Context Arena
Accuracy (8 needles, 4K-128K context)
Memory & Persistence
93.98Thinking Level · Medium
79.53Thinking Level · High
92.17Thinking Level · High
95.95Thinking Level · High
69.42Thinking Enabled
OSWorld-Verified
Accuracy
Desktop Workflows
84.30Thinking Enabled | Tools
81.20Thinking Level · Extra High | Tools
--
--
--
20 additional benchmarks remain in the chart above.

Standard API Pricing: Qwen3.8-27B vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Claude Sonnet 5
Anthropic$2 / 1M tokens$10 / 1M tokens—
GPT-5.6 Terra
OpenAI$2 / 1M tokens$12 / 1M tokens—
Gemini 3.7 Flash
Google DeepMind$0.75 / 1M tokens$3.75 / 1M tokens—
DeepSeek-V4-Flash
DeepSeek-AI$0.14 / 1M tokens$0.28 / 1M tokens—

Version History

How each version of the Qwen3.8-27B series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkQwen3.8-27BCurrentQwen3.6-27BQwen3.5-27BQwen3-32BQwen2.5-32B
HLE
Accuracy
Knowledge Exams
30.80Thinking Enabled
24.00Thinking Enabled
48.50Thinking Enabled | Tools
--
--
CritPt
Score
Scientific Reasoning
5.40Thinking Level · Extra High
1.10Thinking Enabled
0.90Thinking Enabled
0.30Thinking Enabled
--
GPQA Diamond
Accuracy
Scientific Reasoning
89.20Thinking Enabled
87.80Thinking Enabled
85.50Thinking Enabled
68.40Thinking Enabled
--
LiveCodeBench
Pass @K
Algorithmic Coding
90.30Thinking Enabled
83.90Thinking Enabled
80.70Thinking Enabled | Tools
65.70Thinking Enabled
51.20Standard Mode
SWE-Bench Pro - Public
Accuracy
Repository Engineering
61.70Thinking Enabled | Tools
53.50Thinking Enabled | Tools
--
--
--
Context Arena
Accuracy (8 needles, 4K-128K context)
Memory & Persistence
93.98Thinking Level · Medium
82.17Thinking Enabled
73.34Thinking Enabled
--
--
OSWorld-Verified
Accuracy
Desktop Workflows
84.30Thinking Enabled | Tools
--
56.20Thinking Enabled | Tools
--
--
ECI
ECI score (capability index, higher is better)
Capability Indices
149.38Thinking Level · High
146.47Thinking Level · High
--
138.51Thinking Level · High
128.52Thinking Level · High
MMMU-Pro
Accuracy
Visual Understanding
76.30Thinking Level · Extra High
74.60Thinking Enabled
75.00Thinking Enabled
--
--
GDP.pdf
Task pass rate (%)
Documents & Charts
16.40Thinking Level · Extra High
11.00Thinking Enabled
--
1.60Thinking Enabled
--
SciCode
Score
Scientific Computing
46.60Thinking Level · Extra High
42.80Thinking Enabled
--
36.00Thinking Enabled
--
τ³-Banking
Score
Service Workflows
48.00Thinking Level · Extra High | Tools
16.70Thinking Enabled | Tools
--
5.40Thinking Enabled | Tools
--

Single-Benchmark Version Trend

Viewing: HLE · Knowledge Exams

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Qwen3.8-27B Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

Qwen3-32B
Supplier: 阿里巴巴
Standard input: ¥0.0012 / 1K tokens
Standard output: ¥0.0048 / 1K tokens
Qwen2.5-32B
Supplier: 阿里巴巴
Standard input: ¥0.002 / 1K tokens
Standard output: ¥0.006 / 1K tokens
ModelSupplierStandard inputStandard outputBase price applies to
Qwen3-32B
阿里巴巴¥0.0012 / 1K tokens¥0.0048 / 1K tokens—
Qwen2.5-32B
阿里巴巴¥0.002 / 1K tokens¥0.006 / 1K tokens—