DataLearner logo

Gemma 4 E2B Benchmark Details

Gemma 4 E2B currently shows benchmark results led by MMLU-Pro (123 / 176, score 60), LiveCodeBench (180 / 251, score 44), IF Bench (230 / 282, score 38). This page also compares it with 2 competitor models and 1 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

Gemma 4 E2B

Benchmark Results

Thinking
Tool usage

General Knowledge

3 evaluations
Benchmark / mode
Score
Rank/total
MMLU-Pro
Thinking Mode
60
123 / 176
HLE
Standard Mode
4.70
492 / 565
HLE
Thinking Mode
4.80
490 / 565

Other

2 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
40.50
432 / 463
GPQA Diamond
Thinking Mode
43.30
423 / 463

Coding and Software Engineer

2 evaluations
Benchmark / mode
Score
Rank/total
CodeForces
Thinking Mode
633
21 / 21
LiveCodeBench
Thinking Mode
44
180 / 251

Agent Level Benchmark

5 evaluations
Benchmark / mode
Score
Rank/total
τ²-Bench
Thinking Mode
24.50
44 / 44
τ²-Bench - Telecom
Standard ModeTools
22.20
238 / 264
τ²-Bench - Telecom
Thinking ModeTools
20.80
246 / 264
Terminal Bench Hard
Standard ModeTools
2.30
229 / 244
Terminal Bench Hard
Thinking ModeTools
3
225 / 244

Instruction Following

2 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Standard Mode
33.60
250 / 282
IF Bench
Thinking Mode
38
230 / 282

Long Context

1 evaluations
Benchmark / mode
Score
Rank/total
AA-LCR
Standard Mode
16.30
170 / 171

Math and Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
AIME 2026
Thinking Mode
37.50
30 / 30

AI Agent - Tool Usage

1 evaluations
Benchmark / mode
Score
Rank/total
Terminal-Bench 2.1
Thinking ModeTools
0.40
193 / 194

Multimodal Understanding

4 evaluations
Benchmark / mode
Score
Rank/total
MathVision
Thinking Mode
52.40
12 / 12
MMMU-Pro
Standard Mode
41.80
220 / 229
MMMU-Pro
unknown
44.20
214 / 229
MMMU-Pro
Thinking Mode
44.20
214 / 229

Other

1 evaluations
Benchmark / mode
Score
Rank/total
19.10
8 / 8

Competitor Comparison

Benchmark scores for Gemma 4 E2B compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

6 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGemma 4 E2BCurrentMiniCPM5-1B
HLE
Accuracy
综合评估
4.80Thinking Enabled
6.50Thinking Enabled
MMLU-Pro
Accuracy
综合评估
60.00Thinking Enabled
48.85Thinking Enabled
GPQA Diamond
Accuracy
科学与综合推理
43.30Thinking Enabled
26.90Standard Mode
τ²-Bench - Telecom
Accuracy
Agent能力评测
22.20Standard Mode | Tools
82.50Standard Mode | Tools
IF Bench
Accuracy
指令跟随
38.00Thinking Enabled
49.30Thinking Enabled
AA-LCR
Accuracy
长上下文能力
16.30Standard Mode
5.00Standard Mode

Standard API Pricing: Gemma 4 E2B vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

Comparable standard text pricing is not available for these models.

Version History

How each version of the Gemma 4 E2B series stacks up on benchmark tests

Gemma 4 E2BGemma-3n-E2B
No benchmark data matches the selected filters.

Standard API Pricing Across the Gemma 4 E2B Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Gemma-3n-E2B
Google DeepMind$0 / 1M tokens$0 / 1M tokens