DataLearner logo

MiniMax M2.5vsKimi K2.5

Across 9 shared benchmarks, MiniMax M2.5 leads overall: MiniMax M2.5 wins 5, Kimi K2.5 wins 4, with 0 ties and an average score difference of -23.38.

MiniMaxAI
MiniMax M2.5

MiniMaxAI · 2026-02-12 · Reasoning model

Moonshot AI
Kimi K2.5

Moonshot AI · 2026-01-27 · Multimodal model

MiniMax M2.55 wins(56%)(44%)4 winsKimi K2.5

Benchmark scores

Grouped by capability, sorted by largest gap within each. 9 shared benchmarks.

Claw-style Agent Evaluation

MiniMax M2.5 2/2
BenchmarkMiniMax M2.5Kimi K2.5Diff
Claw Bench92.104 / 29Thinking (With Tools)81.7018 / 29Thinking (With Tools)+10.40
Pinch Bench87.807 / 38Thinking (With Tools)84.8018 / 38Thinking (With Tools)+3

Coding and Software Engineer

MiniMax M2.5 2/2
BenchmarkMiniMax M2.5Kimi K2.5Diff
SWE-Bench Pro - Public55.4031 / 62Thinking (With Tools)50.7047 / 62Thinking (With Tools)+4.70
SWE-bench Verified80.2014 / 116Thinking (With Tools)76.8030 / 116Thinking (With Tools)+3.40

AI Agent - Tool Usage

MiniMax M2.5 1/1
BenchmarkMiniMax M2.5Kimi K2.5Diff
Terminal Bench 2.051.7032 / 48Thinking (With Tools)50.8035 / 48Thinking (With Tools)+0.90

General Knowledge

Kimi K2.5 1/1
BenchmarkMiniMax M2.5Kimi K2.5Diff
ARC-AGI-163.6791 / 147Thinking (No Tools)65.3390 / 147Thinking (No Tools)-1.66

Math and Reasoning

Kimi K2.5 1/1
BenchmarkMiniMax M2.5Kimi K2.5Diff
AIME202586.3073 / 215Thinking (No Tools)96.1024 / 215Thinking (No Tools)-9.80

Productivity Knowledge

Kimi K2.5 1/1
BenchmarkMiniMax M2.5Kimi K2.5Diff
GDPval-AA3611 / 15Thinking (No Tools)409 / 15Thinking (No Tools)-4

Writing and Creative Capabilities

Kimi K2.5 1/1
BenchmarkMiniMax M2.5Kimi K2.5Diff
Creative Writing1,35868 / 106Normal (No Tools)1,57643 / 106Normal (No Tools)-217.40

Specs

FieldMiniMax M2.5Kimi K2.5
PublisherMiniMaxAIMoonshot AI
Release date2026-02-122026-01-27
Model typeReasoning modelMultimodal model
ArchitectureMoEMoE
Parameters229B1T
Context length128K256K
Max outputNot available16K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemMiniMax M2.5Kimi K2.5
Text input$0.3 / 1M tokens$0.6 / 1M tokens
Text output$2.4 / 1M tokens$3 / 1M tokens
Cache readNot public$0.1 / 1M tokens

Summary

  • MiniMax M2.5leads in:Claw-style Agent Evaluation (2/2), Coding and Software Engineer (2/2), AI Agent - Tool Usage (1/1)
  • Kimi K2.5leads in:General Knowledge (1/1), Math and Reasoning (1/1), Productivity Knowledge (1/1), Writing and Creative Capabilities (1/1)

On average across the 9 shared benchmarks, Kimi K2.5 scores 23.38 higher.

Largest single-benchmark gap: Creative Writing — MiniMax M2.5 1,358 vs Kimi K2.5 1,576 (-217.40).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.