DataLearner logo

Qwen3-235B-A22B-2507 Benchmark Details

Qwen3-235B-A22B-2507 currently shows benchmark results led by SimpleQA (10 / 47, score 54.30), MMLU Pro (46 / 133, score 83), GPQA Diamond (158 / 270, score 77.50). This page also tracks comparisons against 1 predecessor or same-series models. 1 source link is attached for reference.

Benchmark Results

Qwen3-235B-A22B-2507

Benchmark Results

Thinking

General Knowledge

4 evaluations
Benchmark / mode
Score
Rank/total
MMLU Pro
Standard Mode
83
46 / 133
LiveBench
Standard Mode
48.84
95 / 115
ARC-AGI
Standard Mode
11
61 / 68
ARC-AGI-2
Standard Mode
1.30
55 / 62

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
77.50
158 / 270

Common Sense

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleQA
Standard Mode
54.30
10 / 47

Coding and Software Engineer

1 evaluations
Benchmark / mode
Score
Rank/total
LiveCodeBench
Standard Mode
51.80
94 / 127

Math and Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
AIME2025
Standard Mode
70.30
72 / 106

Version History

How each version of the Qwen3-235B-A22B-2507 series stacks up on benchmark tests

Qwen3-235B-A22B-2507Qwen2.5-72B
Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

2 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkQwen3-235B-A22B-2507CurrentQwen2.5-72B
MMLU Pro
综合评估
83.00Standard Mode
58.10Standard Mode
GPQA Diamond
科学与综合推理
77.50Standard Mode
45.90Standard Mode

Single-Benchmark Version Trend

Viewing: MMLU Pro · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Qwen3-235B-A22B-2507 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

Qwen3-235B-A22B-2507
Supplier: 阿里巴巴
Standard input: ¥0.002 / 1K tokens
Standard output: ¥0.008 / 1K tokens
Qwen2.5-72B
Supplier: 阿里巴巴
Standard input: ¥0.004 / 1K tokens
Standard output: ¥0.012 / 1K tokens
ModelSupplierStandard inputStandard outputBase price applies to
Qwen3-235B-A22B-2507
阿里巴巴¥0.002 / 1K tokens¥0.008 / 1K tokens
Qwen2.5-72B
阿里巴巴¥0.004 / 1K tokens¥0.012 / 1K tokens

Sources