DataLearner logo

MiniMax-M2.7vsKimi K2.5

Across 8 shared benchmarks, MiniMax-M2.7 leads overall: MiniMax-M2.7 wins 5, Kimi K2.5 wins 3, with 0 ties and an average score difference of +0.43.

MiniMaxAI
MiniMax-M2.7

MiniMaxAI · 2026-03-18 · Reasoning model

Moonshot AI
Kimi K2.5

Moonshot AI · 2026-01-27 · Multimodal model

MiniMax-M2.75 wins(63%)(38%)3 winsKimi K2.5

Benchmark scores

Grouped by capability, sorted by largest gap within each. 8 shared benchmarks.

General Knowledge

Kimi K2.5 3/3
BenchmarkMiniMax-M2.7Kimi K2.5Diff
HLE2896 / 172Thinking (No Tools)50.2027 / 172Thinking (With Tools)-22.20
LiveBench63.4956 / 115Deep Thinking (No Tools)69.0742 / 115Thinking (No Tools)-5.58
GPQA Diamond8742 / 187Thinking (No Tools)87.6037 / 187Thinking (No Tools)-0.60

Claw-style Agent Evaluation

MiniMax-M2.7 2/2
BenchmarkMiniMax-M2.7Kimi K2.5Diff
Claw Bench91.705 / 29Thinking (With Tools)81.7018 / 29Thinking (With Tools)+10
Pinch Bench87.109 / 37Thinking (With Tools)84.8017 / 37Thinking (With Tools)+2.30

Coding and Software Engineer

MiniMax-M2.7 1/1
BenchmarkMiniMax-M2.7Kimi K2.5Diff
SWE-Bench Pro - Public56.2024 / 54Thinking (With Tools)50.7041 / 54Thinking (With Tools)+5.50

Long Context

MiniMax-M2.7 1/1
BenchmarkMiniMax-M2.7Kimi K2.5Diff
AA-LCR696 / 15Thinking (With Tools)6512 / 15Thinking (No Tools)+4

Productivity Knowledge

MiniMax-M2.7 1/1
BenchmarkMiniMax-M2.7Kimi K2.5Diff
GDPval-AA5013 / 21Thinking (No Tools)4015 / 21Thinking (No Tools)+10

Specs

FieldMiniMax-M2.7Kimi K2.5
PublisherMiniMaxAIMoonshot AI
Release date2026-03-182026-01-27
Model typeReasoning modelMultimodal model
ArchitectureMoEMoE
Parameters229B1T
Context length200K256K
Max output200K16K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemMiniMax-M2.7Kimi K2.5
Text input$0.3 / 1M tokens$0.6 / 1M tokens
Text output$1.2 / 1M tokens$3 / 1M tokens
Cache read$0.06 / 1M tokens$0.1 / 1M tokens
Cache write$0.375 / 1M tokensNot public

Summary

  • MiniMax-M2.7leads in:Claw-style Agent Evaluation (2/2), Coding and Software Engineer (1/1), Long Context (1/1), Productivity Knowledge (1/1)
  • Kimi K2.5leads in:General Knowledge (3/3)

On average across the 8 shared benchmarks, MiniMax-M2.7 scores 0.43 higher.

Largest single-benchmark gap: HLE — MiniMax-M2.7 28 vs Kimi K2.5 50.20 (-22.20).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.