DataLearner logo

Gemma 4 31BvsKimi K2.5

Across 6 shared benchmarks, Kimi K2.5 leads overall: Gemma 4 31B wins 1, Kimi K2.5 wins 5, with 0 ties and an average score difference of -6.01.

DeepMind
Gemma 4 31B

DeepMind · 2026-04-02 · Chat model

Moonshot AI
Kimi K2.5

Moonshot AI · 2026-01-27 · Multimodal model

Gemma 4 31B1 win(17%)(83%)5 winsKimi K2.5

Benchmark scores

Grouped by capability, sorted by largest gap within each. 6 shared benchmarks.

General Knowledge

Kimi K2.5 3/4
BenchmarkGemma 4 31BKimi K2.5Diff
HLE26.5097 / 172Thinking (With Tools + Internet)50.2027 / 172Thinking (With Tools)-23.70
LiveBench61.6262 / 115Normal (No Tools)69.0742 / 115Thinking (No Tools)-7.45
MMLU Pro85.2024 / 132Thinking (No Tools)78.5069 / 132Thinking (No Tools)+6.70
GPQA Diamond84.3058 / 187Thinking (No Tools)87.6037 / 187Thinking (No Tools)-3.30

Coding and Software Engineer

Kimi K2.5 1/1
BenchmarkGemma 4 31BKimi K2.5Diff
LiveCodeBench8030 / 123Thinking (No Tools)8516 / 123Thinking (No Tools)-5

Math and Reasoning

Kimi K2.5 1/1
BenchmarkGemma 4 31BKimi K2.5Diff
AIME 202689.2015 / 18Thinking (No Tools)92.5012 / 18Thinking (No Tools)-3.30

Specs

FieldGemma 4 31BKimi K2.5
PublisherDeepMindMoonshot AI
Release date2026-04-022026-01-27
Model typeChat modelMultimodal model
ArchitectureDenseMoE
Parameters30.7B1T
Context length256K256K
Max output32K16K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGemma 4 31BKimi K2.5
Text inputNot public$0.6 / 1M tokens
Text outputNot public$3 / 1M tokens
Cache readNot public$0.1 / 1M tokens

One or both models have incomplete public pricing.

Summary

  • Kimi K2.5leads in:General Knowledge (3/4), Coding and Software Engineer (1/1), Math and Reasoning (1/1)

On average across the 6 shared benchmarks, Kimi K2.5 scores 6.01 higher.

Largest single-benchmark gap: HLE — Gemma 4 31B 26.50 vs Kimi K2.5 50.20 (-23.70).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.