Gemma 4 31B Benchmark Details
Gemma 4 31B currently shows benchmark results led by MMLU-Pro (26 / 175, score 85.20), LiveCodeBench (33 / 125, score 80), Terminal Bench Hard (67 / 244, score 36.40). This page also compares it with 3 competitor models and 2 predecessor or same-series models, including performance and pricing views when available. 1 source link is attached for reference.
Benchmark Results
Benchmark Results
Knowledge Exams
3 evaluationsScientific Reasoning
2 evaluationsAlgorithmic Coding
2 evaluationsService Workflows
3 evaluationsAgentic Development
2 evaluationsVisual Understanding
4 evaluationsCompetitor Comparison
Benchmark scores for Gemma 4 31B compared against top models in its class
12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.
| Benchmark | Gemma 4 31BCurrent | GLM-5 | Kimi K2.5 | Qwen3.5-27B |
|---|---|---|---|---|
26.50Thinking Enabled | Tools | 50.40Thinking Enabled | Tools | 50.20Thinking Enabled | Tools | 48.50Thinking Enabled | Tools | |
85.20Thinking Enabled | -- | 78.50Thinking Enabled | 86.10Thinking Enabled | |
1.40Thinking Enabled | 2.00Thinking Enabled | 3.10Thinking Enabled | 0.90Thinking Enabled | |
84.30Thinking Enabled | 86.00Thinking Enabled | 87.60Thinking Enabled | 85.50Thinking Enabled | |
2150.00Thinking Enabled | -- | -- | 1899.00Thinking Enabled | |
80.00Thinking Enabled | -- | 85.00Thinking Enabled | 80.70Thinking Enabled | Tools | |
1368.20Standard Mode | 1600.60Standard Mode | 1578.60Standard Mode | -- | |
76.90Thinking Enabled | Tools | 89.70Thinking Enabled | Tools | -- | 79.00Thinking Enabled | Tools | |
14.80Thinking Enabled | Tools | 9.79Thinking Enabled | Tools | 14.20Thinking Enabled | Tools | -- | |
61.62Standard Mode | 68.85Standard Mode | 69.07Thinking Enabled | -- | |
36.40Thinking Enabled | Tools | 43.00Thinking Enabled | Tools | 34.80Thinking Enabled | Tools | 32.60Thinking Enabled | Tools | |
89.20Thinking Enabled | 92.70Thinking Enabled | 92.50Thinking Enabled | -- |
Standard API Pricing: Gemma 4 31B vs. Peer Models
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
GLM-5 | 智谱AI | $1 / 1M tokens | $3.2 / 1M tokens | — |
Kimi K2.5 | Moonshot AI | $0.6 / 1M tokens | $3 / 1M tokens | — |
Version History
How each version of the Gemma 4 31B series stacks up on benchmark tests
4 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.
| Benchmark | Gemma 4 31BCurrent | Gemma 3 - 27B (IT) | Gemma2-27B |
|---|---|---|---|
85.20Thinking Enabled | 67.50Standard Mode | 56.54Standard Mode | |
84.30Thinking Enabled | 42.40Standard Mode | -- | |
80.00Thinking Enabled | 29.70Standard Mode | -- | |
1368.20Standard Mode | 1265.70Standard Mode | -- |
Single-Benchmark Version Trend
Viewing: MMLU-Pro · Knowledge Exams
Standard API Pricing Across the Gemma 4 31B Series
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
Gemma 3 - 27B (IT) | DeepInfra | $0.09 / 1M tokens | $0.16 / 1M tokens | — |


