DataLearner logo

Gemma 4 31BvsKimi K2.5

Across 11 shared benchmarks, Kimi K2.5 leads overall: Gemma 4 31B wins 3, Kimi K2.5 wins 8, with 0 ties and an average score difference of -2.32.

DeepMind
Gemma 4 31B

DeepMind · 2026-04-02 · Chat model

Moonshot AI
Kimi K2.5

Moonshot AI · 2026-01-27 · Multimodal model

Gemma 4 31B3 wins(27%)(73%)8 winsKimi K2.5

Benchmark scores

Grouped by capability, sorted by largest gap within each. 11 shared benchmarks.

Agent Level Benchmark

Even 2/2
BenchmarkGemma 4 31BKimi K2.5Diff
τ²-Bench - Telecom65.50149 / 264Normal (With Tools)81.30108 / 264Normal (With Tools)-15.80
Terminal Bench Hard30.30110 / 244Normal (With Tools)18.90150 / 244Normal (With Tools)+11.40

General Knowledge

Even 2/2
BenchmarkGemma 4 31BKimi K2.5Diff
MMLU Pro85.2026 / 176Thinking (No Tools)78.5071 / 176Thinking (No Tools)+6.70
HLE11.80369 / 563Normal (No Tools) · Text only13.20357 / 563Normal (No Tools) · Text only-1.40

Claw-style Agent Evaluation

Kimi K2.5 1/1
BenchmarkGemma 4 31BKimi K2.5Diff
PinchBench v252.6638 / 45Reported best (effort unspecified)54.6036 / 45Reported best (effort unspecified)-1.94

Coding and Software Engineer

Kimi K2.5 1/1
BenchmarkGemma 4 31BKimi K2.5Diff
LiveCodeBench8052 / 250Thinking (No Tools)8529 / 250Thinking (No Tools)-5

General Evaluation

Kimi K2.5 1/1
BenchmarkGemma 4 31BKimi K2.5Diff
GPQA Diamond76.30270 / 462Normal (No Tools)78.90248 / 462Normal (No Tools)-2.60

Instruction Following

Gemma 4 31B 1/1
BenchmarkGemma 4 31BKimi K2.5Diff
IF Bench53.50144 / 282Normal (No Tools)43.70188 / 282Normal (No Tools)+9.80

Long Context

Kimi K2.5 1/1
BenchmarkGemma 4 31BKimi K2.5Diff
AA-LCR46.70141 / 170Normal (No Tools)67.30113 / 170Normal (No Tools)-20.60

Math and Reasoning

Kimi K2.5 1/1
BenchmarkGemma 4 31BKimi K2.5Diff
AIME 202689.2025 / 29Thinking (No Tools)92.5021 / 29Thinking (No Tools)-3.30

Multimodal Understanding

Kimi K2.5 1/1
BenchmarkGemma 4 31BKimi K2.5Diff
MMMU-Pro70.30126 / 227Normal (No Tools)73.10110 / 227Normal (No Tools)-2.80

Specs

FieldGemma 4 31BKimi K2.5
PublisherDeepMindMoonshot AI
Release date2026-04-022026-01-27
Model typeChat modelMultimodal model
ArchitectureDenseMoE
Parameters30.7B1T
Context length256K256K
Max output32K16K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGemma 4 31BKimi K2.5
Text inputNot public$0.6 / 1M tokens
Text outputNot public$3 / 1M tokens
Cache readNot public$0.1 / 1M tokens

One or both models have incomplete public pricing.

Summary

  • Gemma 4 31Bleads in:Instruction Following (1/1)
  • Kimi K2.5leads in:Claw-style Agent Evaluation (1/1), Coding and Software Engineer (1/1), General Evaluation (1/1), Long Context (1/1), Math and Reasoning (1/1), Multimodal Understanding (1/1)
  • Tied in:Agent Level Benchmark, General Knowledge

On average across the 11 shared benchmarks, Kimi K2.5 scores 2.32 higher.

Largest single-benchmark gap: AA-LCR — Gemma 4 31B 46.70 vs Kimi K2.5 67.30 (-20.60).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.