DataLearner logo

Mistral Medium 3.5vsGemma 4 31B

Across 3 shared benchmarks, Gemma 4 31B leads overall: Mistral Medium 3.5 wins 1, Gemma 4 31B wins 2, with 0 ties and an average score difference of +4.25.

MistralAI
Mistral Medium 3.5

MistralAI · 2026-05-01 · Chat model

DeepMind
Gemma 4 31B

DeepMind · 2026-04-02 · Chat model

Mistral Medium 3.51 win(33%)(67%)2 winsGemma 4 31B

Benchmark scores

Grouped by capability, sorted by largest gap within each. 3 shared benchmarks.

Coding and Software Engineer

Gemma 4 31B 1/1
BenchmarkMistral Medium 3.5Gemma 4 31BDiff
SciCode39.58104 / 130Thinking (No Tools)45.5087 / 130Thinking (No Tools)-5.92

Multimodal Understanding

Gemma 4 31B 1/1
BenchmarkMistral Medium 3.5Gemma 4 31BDiff
GDP.pdf2.80102 / 118Thinking (No Tools)691 / 118Thinking (No Tools)-3.20

Productivity Knowledge

Mistral Medium 3.5 1/1
BenchmarkMistral Medium 3.5Gemma 4 31BDiff
Harvey Lab-AA69.1034 / 43Thinking (With Tools)47.2340 / 43Thinking (With Tools)+21.87

Specs

FieldMistral Medium 3.5Gemma 4 31B
PublisherMistralAIDeepMind
Release date2026-05-012026-04-02
Model typeChat modelChat model
ArchitectureDenseDense
ParametersNot available30.7B
Context lengthNot available256K
Max outputNot available32K

Summary

  • Mistral Medium 3.5leads in:Productivity Knowledge (1/1)
  • Gemma 4 31Bleads in:Coding and Software Engineer (1/1), Multimodal Understanding (1/1)

On average across the 3 shared benchmarks, Mistral Medium 3.5 scores 4.25 higher.

Largest single-benchmark gap: Harvey Lab-AA — Mistral Medium 3.5 69.10 vs Gemma 4 31B 47.23 (+21.87).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.