DataLearner logo

MiniMax-M2.7vsKimi K2.5

Across 8 shared benchmarks, MiniMax-M2.7 leads overall: MiniMax-M2.7 wins 7, Kimi K2.5 wins 1, with 0 ties and an average score difference of +6.44.

MiniMaxAI
MiniMax-M2.7

MiniMaxAI · 2026-03-18 · Reasoning model

Moonshot AI
Kimi K2.5

Moonshot AI · 2026-01-27 · Multimodal model

MiniMax-M2.77 wins(88%)(13%)1 winKimi K2.5

Benchmark scores

Grouped by capability, sorted by largest gap within each. 8 shared benchmarks.

Claw-style Agent Evaluation

MiniMax-M2.7 3/3
BenchmarkMiniMax-M2.7Kimi K2.5Diff
PinchBench v266.7530 / 45Reported best (effort unspecified)54.6036 / 45Reported best (effort unspecified)+12.15
Claw Bench91.705 / 29Thinking (With Tools)81.7018 / 29Thinking (With Tools)+10
Pinch Bench87.1010 / 38Thinking (With Tools)84.8018 / 38Thinking (With Tools)+2.30

AI Agent - Tool Usage

MiniMax-M2.7 2/2
BenchmarkMiniMax-M2.7Kimi K2.5Diff
Terminal-Bench 2.155.40121 / 192Thinking (With Tools)45.70137 / 192Thinking (With Tools)+9.70
Terminal Bench 2.05725 / 48Thinking (With Tools)50.8035 / 48Thinking (With Tools)+6.20

Agent Level Benchmark

Kimi K2.5 1/1
BenchmarkMiniMax-M2.7Kimi K2.5Diff
τ³-Banking9.90129 / 164Thinking (With Tools)14.20113 / 164Thinking (With Tools)-4.30

Coding and Software Engineer

MiniMax-M2.7 1/1
BenchmarkMiniMax-M2.7Kimi K2.5Diff
SWE-Bench Pro - Public56.2029 / 62Thinking (With Tools)50.7047 / 62Thinking (With Tools)+5.50

Productivity Knowledge

MiniMax-M2.7 1/1
BenchmarkMiniMax-M2.7Kimi K2.5Diff
GDPval-AA507 / 15Thinking (No Tools)409 / 15Thinking (No Tools)+10

Specs

FieldMiniMax-M2.7Kimi K2.5
PublisherMiniMaxAIMoonshot AI
Release date2026-03-182026-01-27
Model typeReasoning modelMultimodal model
ArchitectureMoEMoE
Parameters229B1T
Context length200K256K
Max output200K16K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemMiniMax-M2.7Kimi K2.5
Text input$0.3 / 1M tokens$0.6 / 1M tokens
Text output$1.2 / 1M tokens$3 / 1M tokens
Cache read$0.06 / 1M tokens$0.1 / 1M tokens
Cache write$0.375 / 1M tokensNot public

Summary

  • MiniMax-M2.7leads in:Claw-style Agent Evaluation (3/3), AI Agent - Tool Usage (2/2), Coding and Software Engineer (1/1), Productivity Knowledge (1/1)
  • Kimi K2.5leads in:Agent Level Benchmark (1/1)

On average across the 8 shared benchmarks, MiniMax-M2.7 scores 6.44 higher.

Largest single-benchmark gap: PinchBench v2 — MiniMax-M2.7 66.75 vs Kimi K2.5 54.60 (+12.15).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.