DataLearner logo

Qwen3.6-35B-A3B Benchmark Details

Qwen3.6-35B-A3B currently shows benchmark results led by GPQA (2 / 17, score 86), C-Eval (7 / 48, score 90), MMLU Pro (25 / 134, score 85.20). This page also compares it with 3 competitor models and 1 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

Qwen3.6-35B-A3B

Benchmark Results

Thinking

General Knowledge

4 evaluations
Benchmark / mode
Score
Rank/total
C-Eval
Thinking Mode
90
7 / 48
GPQA
Thinking Mode
86
2 / 17
MMLU Pro
Thinking Mode
85.20
25 / 134
HLE
Thinking Mode
21.40
137 / 197

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
84.85
103 / 274

Coding and Software Engineer

4 evaluations
Benchmark / mode
Score
Rank/total
LiveCodeBench
Thinking Mode
80.40
31 / 128
SWE-bench Verified
Thinking Mode
73.40
48 / 116
67.20
26 / 29
49.50
49 / 62

AI Agent - Tool Usage

2 evaluations
Benchmark / mode
Score
Rank/total
Terminal Bench 2.0
Thinking Mode
51.50
33 / 48
Tool Decathlon
Thinking Mode
26.90
10 / 10

Text Embedding

2 evaluations
Benchmark / mode
Score
Rank/total
Context Arena
Standard Mode
53.20
84 / 126
Context Arena
Thinking Mode
83.53
27 / 126

Math and Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
AIME 2026
Thinking Mode
92.70
10 / 21
IMO-AnswerBench
Thinking Mode
78.90
21 / 24

Competitor Comparison

Benchmark scores for Qwen3.6-35B-A3B compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

7 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkQwen3.6-35B-A3BCurrentGemma 4 26B A4BGLM-4.7-FlashMistral Medium 3.5
HLE
Accuracy
综合评估
21.40Thinking Enabled
17.20Thinking Enabled | Tools
14.40Thinking Enabled
--
MMLU Pro
Accuracy
综合评估
85.20Thinking Enabled
82.60Thinking Enabled
--
--
GPQA Diamond
Accuracy
科学与综合推理
84.85Standard Mode
82.30Thinking Enabled
75.20Thinking Enabled
--
LiveCodeBench
Pass @K
编程与软件工程
80.40Thinking Enabled
77.10Thinking Enabled
--
--
SWE-bench Verified
Accuracy
编程与软件工程
73.40Thinking Enabled
--
59.20Thinking Enabled
--
Context Arena
Accuracy (8 needles, 4K-128K context)
文本向量检索
83.53Thinking Enabled
--
--
32.05Thinking Level · High
AIME 2026
Accuracy
数学推理
92.70Thinking Enabled
88.30Thinking Enabled
--
--

Standard API Pricing: Qwen3.6-35B-A3B vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · CNY / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
GLM-4.7-Flash
智谱AI¥0 / 1M tokens¥0 / 1M tokens

Version History

How each version of the Qwen3.6-35B-A3B series stacks up on benchmark tests

Qwen3.6-35B-A3BQwen3.5-35B-A3B
Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

1 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkQwen3.6-35B-A3BCurrentQwen3.5-35B-A3B
Context Arena
Accuracy (8 needles, 4K-128K context)
文本向量检索
83.53Thinking Enabled
67.48Thinking Enabled

Single-Benchmark Version Trend

Viewing: Context Arena · 文本向量检索

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Qwen3.6-35B-A3B Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

Comparable standard text pricing is not available for these models.