DataLearner logo

MiniMax M2.5vsMiniMax M2

Across 9 shared benchmarks, MiniMax M2.5 leads overall: MiniMax M2.5 wins 8, MiniMax M2 wins 1, with 0 ties and an average score difference of +9.41.

MiniMaxAI
MiniMax M2.5

MiniMaxAI · 2026-02-12 · Reasoning model

MiniMaxAI
MiniMax M2

MiniMaxAI · 2025-10-27 · Chat model

MiniMax M2.58 wins(89%)(11%)1 winMiniMax M2

Benchmark scores

Grouped by capability, sorted by largest gap within each. 9 shared benchmarks.

Agent Level Benchmark

MiniMax M2.5 2/2
BenchmarkMiniMax M2.5MiniMax M2Diff
τ²-Bench - Telecom97.8012 / 264Thinking (With Tools)8777 / 264Thinking (With Tools)+10.80
Terminal Bench Hard34.8075 / 244Thinking (With Tools)25.80126 / 244Thinking (With Tools)+9

General Knowledge

MiniMax M2.5 2/2
BenchmarkMiniMax M2.5MiniMax M2Diff
HLE20.50281 / 563Thinking (No Tools) · Text only13.70351 / 563Thinking (No Tools) · Text only+6.80
CritPt1.10147 / 200Thinking (No Tools)0.90158 / 200Thinking (No Tools)+0.20

AI Agent - Information Search

MiniMax M2.5 1/1
BenchmarkMiniMax M2.5MiniMax M2Diff
BrowseComp76.3025 / 57Thinking (With Tools)4449 / 57Thinking (With Tools)+32.30

Coding and Software Engineer

MiniMax M2.5 1/1
BenchmarkMiniMax M2.5MiniMax M2Diff
SWE-bench Verified80.2014 / 116Thinking (With Tools)69.4064 / 116Thinking (With Tools)+10.80

General Evaluation

MiniMax M2.5 1/1
BenchmarkMiniMax M2.5MiniMax M2Diff
GPQA Diamond85.20157 / 462Thinking (No Tools)78256 / 462Thinking (No Tools)+7.20

Instruction Following

MiniMax M2 1/1
BenchmarkMiniMax M2.5MiniMax M2Diff
IF Bench71.6055 / 282Thinking (No Tools)72.3052 / 282Thinking (No Tools)-0.70

Math and Reasoning

MiniMax M2.5 1/1
BenchmarkMiniMax M2.5MiniMax M2Diff
AIME202586.3073 / 215Thinking (No Tools)78101 / 215Thinking (No Tools)+8.30

Specs

FieldMiniMax M2.5MiniMax M2
PublisherMiniMaxAIMiniMaxAI
Release date2026-02-122025-10-27
Model typeReasoning modelChat model
ArchitectureMoEMoE
Parameters229B230B
Context length128K205K
Max outputNot availableNot available

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemMiniMax M2.5MiniMax M2
Text input$0.3 / 1M tokens¥2.1 / 1M tokens
Text output$2.4 / 1M tokens¥8.4 / 1M tokens
Cache readNot public¥0.21 / 1M tokens
Cache writeNot public¥2.625 / 1M tokens

Summary

  • MiniMax M2.5leads in:Agent Level Benchmark (2/2), General Knowledge (2/2), AI Agent - Information Search (1/1), Coding and Software Engineer (1/1), General Evaluation (1/1), Math and Reasoning (1/1)
  • MiniMax M2leads in:Instruction Following (1/1)

On average across the 9 shared benchmarks, MiniMax M2.5 scores 9.41 higher.

Largest single-benchmark gap: BrowseComp — MiniMax M2.5 76.30 vs MiniMax M2 44 (+32.30).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.