DataLearner logo

Gemma 4 26B A4B Benchmark Details

Gemma 4 26B A4B currently shows benchmark results led by MMLU-Pro (51 / 175, score 82.60), LiveCodeBench (37 / 125, score 77.10), GPQA Diamond (114 / 253, score 82.30). This page also compares it with 2 competitor models and 1 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

Gemma 4 26B A4B

Benchmark Results

Thinking
Tool usage
Internet

Knowledge Exams

3 evaluations
Benchmark / mode
Score
Rank/total
MMLU-Pro
Thinking Mode
82.60
51 / 175
HLE
Thinking Mode
8.70
201 / 235
HLE
Thinking ModeToolsInternet
17.20
174 / 235

Scientific Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Mode
82.30
114 / 253

Algorithmic Coding

2 evaluations
Benchmark / mode
Score
Rank/total
CodeForces
Thinking Mode
1718
18 / 21
LiveCodeBench
Thinking Mode
77.10
37 / 125

Writing

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1304.60
77 / 110

Service Workflows

2 evaluations
Benchmark / mode
Score
Rank/total
τ²-Bench
Thinking ModeTools
68.20
27 / 44
τ³-Banking
Thinking ModeTools
12
126 / 167

Agentic Development

2 evaluations
Benchmark / mode
Score
Rank/total
Terminal Bench Hard
Standard ModeTools
25
128 / 244
Terminal Bench Hard
Thinking ModeTools
13.60
173 / 244

Mathematics

1 evaluations
Benchmark / mode
Score
Rank/total
AIME 2026
Thinking Mode
88.30
27 / 30

Visual Understanding

4 evaluations
Benchmark / mode
Score
Rank/total
MathVision
Thinking Mode
82.40
10 / 12
MMMU-Pro
Standard Mode
66.70
148 / 229
MMMU-Pro
unknown
73.80
104 / 229
MMMU-Pro
Thinking Mode
69.20
134 / 229

Long Retrieval

1 evaluations
Benchmark / mode
Score
Rank/total
44.10
6 / 8

Tool Orchestration

1 evaluations
Benchmark / mode
Score
Rank/total
56.37
34 / 45

Competitor Comparison

Benchmark scores for Gemma 4 26B A4B compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

10 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGemma 4 26B A4BCurrentQwen3.6-35B-A3BGLM-4.7-Flash
HLE
Accuracy
Knowledge Exams
17.20Thinking Enabled | Tools
21.40Thinking Enabled
14.40Thinking Enabled
MMLU-Pro
Accuracy
Knowledge Exams
82.60Thinking Enabled
85.20Thinking Enabled
--
GPQA Diamond
Accuracy
Scientific Reasoning
82.30Thinking Enabled
84.85Standard Mode
75.20Thinking Enabled
LiveCodeBench
Pass @K
Algorithmic Coding
77.10Thinking Enabled
80.40Thinking Enabled
--
Creative Writing
Elo、大模型评判两两对战
Writing
1304.60Standard Mode
--
1124.70Standard Mode
τ²-Bench
Accuracy
Service Workflows
68.20Thinking Enabled | Tools
--
79.50Thinking Enabled | Tools
τ³-Banking
Score
Service Workflows
12.00Thinking Enabled | Tools
9.30Thinking Enabled | Tools
--
Terminal Bench Hard
Accuracy
Agentic Development
25.00Standard Mode | Tools
34.80Thinking Enabled | Tools
32.00Thinking Enabled | Tools
AIME 2026
Accuracy
Mathematics
88.30Thinking Enabled
92.70Thinking Enabled
--
MMMU-Pro
Accuracy
Visual Understanding
73.80Thinking Level · High
75.00Thinking Enabled
--

Standard API Pricing: Gemma 4 26B A4B vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · CNY / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
GLM-4.7-Flash
智谱AI¥0 / 1M tokens¥0 / 1M tokens—

Version History

How each version of the Gemma 4 26B A4B series stacks up on benchmark tests

Gemma 4 26B A4BGemma 3 - 27B (IT)
Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

4 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGemma 4 26B A4BCurrentGemma 3 - 27B (IT)
MMLU-Pro
Accuracy
Knowledge Exams
82.60Thinking Enabled
67.50Standard Mode
GPQA Diamond
Accuracy
Scientific Reasoning
82.30Thinking Enabled
42.40Standard Mode
LiveCodeBench
Pass @K
Algorithmic Coding
77.10Thinking Enabled
29.70Standard Mode
Creative Writing
Elo、大模型评判两两对战
Writing
1304.60Standard Mode
1265.70Standard Mode

Single-Benchmark Version Trend

Viewing: MMLU-Pro · Knowledge Exams

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Gemma 4 26B A4B Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Gemma 3 - 27B (IT)
DeepInfra$0.09 / 1M tokens$0.16 / 1M tokens—