DataLearner logo

Gemma 4 31B Benchmark Details

Gemma 4 31B currently shows benchmark results led by MMLU-Pro (26 / 175, score 85.20), LiveCodeBench (33 / 125, score 80), Terminal Bench Hard (67 / 244, score 36.40). This page also compares it with 3 competitor models and 2 predecessor or same-series models, including performance and pricing views when available. 1 source link is attached for reference.

Benchmark Results

Gemma 4 31B

Benchmark Results

Thinking
Tool usage
Internet

Knowledge Exams

3 evaluations
Benchmark / mode
Score
Rank/total
MMLU-Pro
Thinking Mode
85.20
26 / 175
HLE
Thinking Mode
19.50
158 / 235
HLE
Thinking ModeToolsInternet
26.50
129 / 235

Scientific Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Mode
84.30
96 / 253
CritPt
Thinking Mode
1.40
139 / 204

Algorithmic Coding

2 evaluations
Benchmark / mode
Score
Rank/total
CodeForces
Thinking Mode
2150
14 / 21
LiveCodeBench
Thinking Mode
80
33 / 125

Writing

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1368.20
70 / 110

Service Workflows

3 evaluations
Benchmark / mode
Score
Rank/total
τ²-Bench
Thinking ModeTools
76.90
20 / 44
τ³-Banking
Standard ModeTools
8.90
140 / 167
τ³-Banking
Thinking ModeTools
14.80
113 / 167

Cross-capability Suites

1 evaluations
Benchmark / mode
Score
Rank/total
LiveBench
Standard Mode
61.62
64 / 117

Agentic Development

2 evaluations
Benchmark / mode
Score
Rank/total
Terminal Bench Hard
Standard ModeTools
30.30
110 / 244
Terminal Bench Hard
Thinking ModeTools
36.40
67 / 244

Mathematics

1 evaluations
Benchmark / mode
Score
Rank/total
AIME 2026
Thinking Mode
89.20
26 / 30

Visual Understanding

4 evaluations
Benchmark / mode
Score
Rank/total
MathVision
Thinking Mode
85.60
9 / 12
MMMU-Pro
Standard Mode
70.30
128 / 229
MMMU-Pro
unknown
76.90
71 / 229
MMMU-Pro
Thinking Mode
73.40
109 / 229

Long Retrieval

1 evaluations
Benchmark / mode
Score
Rank/total
66.40
5 / 8

Scientific Computing

1 evaluations
Benchmark / mode
Score
Rank/total
SciCode
Thinking Mode
45.50
91 / 134

Legal

1 evaluations
Benchmark / mode
Score
Rank/total
Harvey Lab-AA
Thinking ModeTools
47.23
40 / 44

Documents & Charts

1 evaluations
Benchmark / mode
Score
Rank/total
GDP.pdf
Thinking Mode
6
95 / 122

Tool Orchestration

1 evaluations
Benchmark / mode
Score
Rank/total
52.66
38 / 45

Competitor Comparison

Benchmark scores for Gemma 4 31B compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGemma 4 31BCurrentGLM-5Kimi K2.5Qwen3.5-27B
HLE
Accuracy
Knowledge Exams
26.50Thinking Enabled | Tools
50.40Thinking Enabled | Tools
50.20Thinking Enabled | Tools
48.50Thinking Enabled | Tools
MMLU-Pro
Accuracy
Knowledge Exams
85.20Thinking Enabled
--
78.50Thinking Enabled
86.10Thinking Enabled
CritPt
Score
Scientific Reasoning
1.40Thinking Enabled
2.00Thinking Enabled
3.10Thinking Enabled
0.90Thinking Enabled
GPQA Diamond
Accuracy
Scientific Reasoning
84.30Thinking Enabled
86.00Thinking Enabled
87.60Thinking Enabled
85.50Thinking Enabled
CodeForces
Accuracy
Algorithmic Coding
2150.00Thinking Enabled
--
--
1899.00Thinking Enabled
LiveCodeBench
Pass @K
Algorithmic Coding
80.00Thinking Enabled
--
85.00Thinking Enabled
80.70Thinking Enabled | Tools
Creative Writing
Elo、大模型评判两两对战
Writing
1368.20Standard Mode
1600.60Standard Mode
1578.60Standard Mode
--
τ²-Bench
Accuracy
Service Workflows
76.90Thinking Enabled | Tools
89.70Thinking Enabled | Tools
--
79.00Thinking Enabled | Tools
τ³-Banking
Score
Service Workflows
14.80Thinking Enabled | Tools
9.79Thinking Enabled | Tools
14.20Thinking Enabled | Tools
--
LiveBench
Accuracy
Cross-capability Suites
61.62Standard Mode
68.85Standard Mode
69.07Thinking Enabled
--
Terminal Bench Hard
Accuracy
Agentic Development
36.40Thinking Enabled | Tools
43.00Thinking Enabled | Tools
34.80Thinking Enabled | Tools
32.60Thinking Enabled | Tools
AIME 2026
Accuracy
Mathematics
89.20Thinking Enabled
92.70Thinking Enabled
92.50Thinking Enabled
--
2 additional benchmarks remain in the chart above.

Standard API Pricing: Gemma 4 31B vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
GLM-5
智谱AI$1 / 1M tokens$3.2 / 1M tokens—
Kimi K2.5
Moonshot AI$0.6 / 1M tokens$3 / 1M tokens—

Version History

How each version of the Gemma 4 31B series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

4 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGemma 4 31BCurrentGemma 3 - 27B (IT)Gemma2-27B
MMLU-Pro
Accuracy
Knowledge Exams
85.20Thinking Enabled
67.50Standard Mode
56.54Standard Mode
GPQA Diamond
Accuracy
Scientific Reasoning
84.30Thinking Enabled
42.40Standard Mode
--
LiveCodeBench
Pass @K
Algorithmic Coding
80.00Thinking Enabled
29.70Standard Mode
--
Creative Writing
Elo、大模型评判两两对战
Writing
1368.20Standard Mode
1265.70Standard Mode
--

Single-Benchmark Version Trend

Viewing: MMLU-Pro · Knowledge Exams

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Gemma 4 31B Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Gemma 3 - 27B (IT)
DeepInfra$0.09 / 1M tokens$0.16 / 1M tokens—

Sources