DataLearner logo

Gemma 4 E4B Benchmark Details

Gemma 4 E4B currently shows benchmark results led by MMLU-Pro (100 / 176, score 69.40), LiveCodeBench (159 / 251, score 52), IF Bench (185 / 282, score 44.20). This page also compares it with 2 competitor models and 1 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

Gemma 4 E4B

Benchmark Results

Thinking
Tool usage

General Knowledge

5 evaluations
Benchmark / mode
Score
Rank/total
MMLU-Pro
Thinking Mode
69.40
100 / 176
HLE
Standard Mode
4.80
490 / 565
HLE
Thinking Mode
3.80
535 / 565
CritPt
Standard Mode
0.30
183 / 201
CritPt
Thinking Mode
0.60
169 / 201

Other

2 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
54.90
400 / 463
GPQA Diamond
Thinking Mode
58.60
387 / 463

Coding and Software Engineer

2 evaluations
Benchmark / mode
Score
Rank/total
CodeForces
Thinking Mode
940
20 / 21
LiveCodeBench
Thinking Mode
52
159 / 251

Agent Level Benchmark

5 evaluations
Benchmark / mode
Score
Rank/total
τ²-Bench
Thinking Mode
42.20
38 / 44
τ²-Bench - Telecom
Standard ModeTools
26
225 / 264
τ²-Bench - Telecom
Thinking ModeTools
20.80
246 / 264
Terminal Bench Hard
Standard ModeTools
7.60
190 / 244
Terminal Bench Hard
Thinking ModeTools
8.30
186 / 244

Instruction Following

2 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Standard Mode
40.50
214 / 282
IF Bench
Thinking Mode
44.20
185 / 282

Long Context

1 evaluations
Benchmark / mode
Score
Rank/total
AA-LCR
Standard Mode
24
166 / 171

Math and Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
AIME 2026
Thinking Mode
42.50
29 / 30

AI Agent - Tool Usage

1 evaluations
Benchmark / mode
Score
Rank/total
Terminal-Bench 2.1
Thinking ModeTools
1.90
193 / 196

Multimodal Understanding

4 evaluations
Benchmark / mode
Score
Rank/total
MathVision
Thinking Mode
59.50
11 / 12
MMMU-Pro
Standard Mode
51.20
199 / 229
MMMU-Pro
unknown
52.60
195 / 229
MMMU-Pro
Thinking Mode
52.60
195 / 229

Other

1 evaluations
Benchmark / mode
Score
Rank/total
25.40
7 / 8

Competitor Comparison

Benchmark scores for Gemma 4 E4B compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

6 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGemma 4 E4BCurrentMiniCPM5-1B
HLE
Accuracy
综合评估
4.80Standard Mode
6.50Thinking Enabled
MMLU-Pro
Accuracy
综合评估
69.40Thinking Enabled
48.85Thinking Enabled
GPQA Diamond
Accuracy
科学与综合推理
58.60Thinking Enabled
26.90Standard Mode
τ²-Bench - Telecom
Accuracy
Agent能力评测
26.00Standard Mode | Tools
82.50Standard Mode | Tools
IF Bench
Accuracy
指令跟随
44.20Thinking Enabled
49.30Thinking Enabled
AA-LCR
Accuracy
长上下文能力
24.00Standard Mode
5.00Standard Mode

Standard API Pricing: Gemma 4 E4B vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

Comparable standard text pricing is not available for these models.

Version History

How each version of the Gemma 4 E4B series stacks up on benchmark tests

Gemma 4 E4BGemma-3n-E4B
No benchmark data matches the selected filters.

Standard API Pricing Across the Gemma 4 E4B Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Gemma-3n-E4B
Google DeepMind$0 / 1M tokens$0 / 1M tokens