DataLearner logo

Gemma 4 31B Benchmark Details

Gemma 4 31B currently shows benchmark results led by MMLU Pro (25 / 133, score 85.20), LiveCodeBench (32 / 126, score 80), GPQA Diamond (93 / 226, score 84.30). This page also compares it with 3 competitor models and 2 predecessor or same-series models, including performance and pricing views when available. 1 source link is attached for reference.

Benchmark Results

Gemma 4 31B

Benchmark Results

Thinking
Tool usage
Internet

General Knowledge

4 evaluations
Benchmark / mode
Score
Rank/total
MMLU Pro
Thinking Mode
85.20
25 / 133
LiveBench
Standard Mode
61.62
62 / 115
HLE
Thinking Mode
19.50
129 / 181
HLE
Thinking ModeToolsInternet
26.50
104 / 181

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Mode
84.30
93 / 226

Coding and Software Engineer

2 evaluations
Benchmark / mode
Score
Rank/total
CodeForces
Thinking Mode
2150
13 / 20
LiveCodeBench
Thinking Mode
80
32 / 126

Agent Level Benchmark

1 evaluations
Benchmark / mode
Score
Rank/total
τ²-Bench
Thinking ModeTools
76.90
20 / 43

Math and Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
AIME 2026
Thinking Mode
89.20
16 / 19

Multimodal Understanding

1 evaluations
Benchmark / mode
Score
Rank/total
MathVision
Thinking Mode
85.60
7 / 10

Other

1 evaluations
Benchmark / mode
Score
Rank/total
66.40
5 / 8

Competitor Comparison

Benchmark scores for Gemma 4 31B compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

8 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGemma 4 31BCurrentGLM-5Kimi K2.5Qwen3.5-27B
HLE
综合评估
26.50Thinking Enabled | Tools
50.40Thinking Enabled | Tools
50.20Thinking Enabled | Tools
48.50Thinking Enabled | Tools
LiveBench
综合评估
61.62Standard Mode
68.85Standard Mode
69.07Thinking Enabled
--
MMLU Pro
综合评估
85.20Thinking Enabled
--
78.50Thinking Enabled
86.10Thinking Enabled
GPQA Diamond
科学与综合推理
84.30Thinking Enabled
86.00Thinking Enabled
87.60Thinking Enabled
85.50Thinking Enabled
CodeForces
编程与软件工程
2150.00Thinking Enabled
--
--
1899.00Thinking Enabled
LiveCodeBench
编程与软件工程
80.00Thinking Enabled
--
85.00Thinking Enabled
80.70Thinking Enabled | Tools
τ²-Bench
Agent能力评测
76.90Thinking Enabled | Tools
89.70Thinking Enabled | Tools
--
79.00Thinking Enabled | Tools
AIME 2026
数学推理
89.20Thinking Enabled
92.70Thinking Enabled
92.50Thinking Enabled
--

Standard API Pricing: Gemma 4 31B vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
GLM-5
智谱AI$1 / 1M tokens$3.2 / 1M tokens
Kimi K2.5
Moonshot AI$0.6 / 1M tokens$3 / 1M tokens

Version History

How each version of the Gemma 4 31B series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

3 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGemma 4 31BCurrentGemma 3 - 27B (IT)Gemma2-27B
MMLU Pro
综合评估
85.20Thinking Enabled
67.50Standard Mode
56.54Standard Mode
GPQA Diamond
科学与综合推理
84.30Thinking Enabled
42.40Standard Mode
--
LiveCodeBench
编程与软件工程
80.00Thinking Enabled
29.70Standard Mode
--

Single-Benchmark Version Trend

Viewing: MMLU Pro · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Gemma 4 31B Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Gemma 3 - 27B (IT)
DeepInfra$0.09 / 1M tokens$0.16 / 1M tokens

Sources