Gemma 4 31B Benchmark Details
Gemma 4 31B currently shows benchmark results led by MMLU Pro (25 / 133, score 85.20), LiveCodeBench (32 / 126, score 80), GPQA Diamond (93 / 226, score 84.30). This page also compares it with 3 competitor models and 2 predecessor or same-series models, including performance and pricing views when available. 1 source link is attached for reference.
Benchmark Results
Benchmark Results
General Knowledge
4 evaluationsCoding and Software Engineer
2 evaluationsCompetitor Comparison
Benchmark scores for Gemma 4 31B compared against top models in its class
8 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.
| Benchmark | Gemma 4 31BCurrent | GLM-5 | Kimi K2.5 | Qwen3.5-27B |
|---|---|---|---|---|
HLE 综合评估 | 26.50Thinking Enabled | Tools | 50.40Thinking Enabled | Tools | 50.20Thinking Enabled | Tools | 48.50Thinking Enabled | Tools |
LiveBench 综合评估 | 61.62Standard Mode | 68.85Standard Mode | 69.07Thinking Enabled | -- |
MMLU Pro 综合评估 | 85.20Thinking Enabled | -- | 78.50Thinking Enabled | 86.10Thinking Enabled |
GPQA Diamond 科学与综合推理 | 84.30Thinking Enabled | 86.00Thinking Enabled | 87.60Thinking Enabled | 85.50Thinking Enabled |
CodeForces 编程与软件工程 | 2150.00Thinking Enabled | -- | -- | 1899.00Thinking Enabled |
LiveCodeBench 编程与软件工程 | 80.00Thinking Enabled | -- | 85.00Thinking Enabled | 80.70Thinking Enabled | Tools |
τ²-Bench Agent能力评测 | 76.90Thinking Enabled | Tools | 89.70Thinking Enabled | Tools | -- | 79.00Thinking Enabled | Tools |
AIME 2026 数学推理 | 89.20Thinking Enabled | 92.70Thinking Enabled | 92.50Thinking Enabled | -- |
Standard API Pricing: Gemma 4 31B vs. Peer Models
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
GLM-5 | 智谱AI | $1 / 1M tokens | $3.2 / 1M tokens | — |
Kimi K2.5 | Moonshot AI | $0.6 / 1M tokens | $3 / 1M tokens | — |
Version History
How each version of the Gemma 4 31B series stacks up on benchmark tests
3 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.
| Benchmark | Gemma 4 31BCurrent | Gemma 3 - 27B (IT) | Gemma2-27B |
|---|---|---|---|
MMLU Pro 综合评估 | 85.20Thinking Enabled | 67.50Standard Mode | 56.54Standard Mode |
GPQA Diamond 科学与综合推理 | 84.30Thinking Enabled | 42.40Standard Mode | -- |
LiveCodeBench 编程与软件工程 | 80.00Thinking Enabled | 29.70Standard Mode | -- |
Single-Benchmark Version Trend
Viewing: MMLU Pro · 综合评估
Standard API Pricing Across the Gemma 4 31B Series
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
Gemma 3 - 27B (IT) | DeepInfra | $0.09 / 1M tokens | $0.16 / 1M tokens | — |