MiniMax M3 leads overall
Ahead on 4 of 5 benchmarks, averaging 6.6 points higher
Summarised from the 5 benchmarks both models were scored on; details in the charts below. A further 1 Elo/rating-scale benchmarks are left out of the average — their scale cannot be added to percentages.

MiniMax-M2.7
MiniMaxAI
Updates live with the mode filters below.
Best overall
MiniMax-M2.7 · 300.66
Best single
MiniMax-M2.7 · GDPval-AA v2 1495.00
Modality coverage
MiniMax M3 · 3 modalities
Head to head
6
Benchmarks
2
Wins
4
Losses
+13.74
Average diff
Benchmark-by-benchmark comparison. Changing the thinking mode or tool filters updates the chart and table below.
Data sourced primarily from official releases (GitHub, Hugging Face, papers), then benchmark leaderboards, then third-party evaluators. Learn about our data methodology
Each axis is the mean percentage score of one benchmark domain. It is an average, not a capability rating.
Relative edge: 科学与综合推理 +5.7 / Relative gap: 文本向量检索 -17.9
Relative edge: 文本向量检索 +17.9 / Relative gap: 科学与综合推理 -5.7
Method: for each model and benchmark, all scores in the current mode scope are averaged (not the best score), then those benchmark scores are averaged within each domain. Only benchmarks scored on a 0-100 scale by at least two of the selected models count — Elo and rating-scale benchmarks such as Codeforces or Arena are excluded, because averaging a 1500 rating with an 85% accuracy produces a meaningless number. Missing values are not counted as zero, and the averages are unweighted, so domains with harder benchmarks read lower.
Every model and runtime mode, benchmark by benchmark. Values are comparable along a row, not between different benchmarks.
6 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.
| Benchmark | MiniMax-M2.7 | MiniMax M3 |
|---|---|---|
LiveBench 综合评估 | 63.49Deep Thinking Mode | 70.02Deep Thinking Mode |
GPQA Diamond 科学与综合推理 | 87.00Thinking Enabled | 81.31Standard Mode |
SWE-Bench Pro - Public 编程与软件工程 | 56.20Thinking Enabled | Tools | 59.00Thinking Enabled | Tools |
Context Arena 文本向量检索 | 33.29Thinking Enabled | 51.15Thinking Enabled |
GDPval-AA v2 生产力知识 | 1495.00Thinking Enabled | Tools | 1379.75Thinking Enabled | Tools |
AA-LCR 长上下文能力 | 69.00Thinking Enabled | Tools | 80.33Thinking Enabled |
Official list prices per model API, split by input and output. Unit: USD per 1M tokens.
Architecture, licensing and API modalities. "Not provided" means the field is missing from our database.
| Features & specs | MiniMax-M2.7MiniMaxAI | MiniMax M3MiniMaxAI |
|---|---|---|
Core specsRelease | 2026-03-18 | 2026-06-01 |
Context length | 200K | 1M |
Total parameters | 229B | 428B |
Active parameters | 10B | 23B |
Max output length | 204,800 tokens | 524,288 tokens |
Architecture | MoE (mixture of experts) | MoE (mixture of experts) |
Runtime modes | 开启关闭 | 开启关闭 |
LicenseCode Open Source | Open Source · MiniMax-Modified MIT | Open Source · MIT License |
Weights Open Source | Open Source · MiniMax-Modified MIT | Open Source · MiniMax-Modified MIT |
Licensing status | 不可以商用 | 不可以商用 |
Local deploymentWeight size | 未知 | Not provided |
VRAM for weights | Not provided | Not provided |
Weights | Hugging Face | Hugging Face |
Source repo | GitHub | GitHub |
API modality supportText Input/Output | / | / |
Image Input/Output | Not provided | / |
Video Input/Output | Not provided | / |
ResourcesPaper / report | MiniMax M2.7: Early Echoes of Self-Evolution | MiniMax M3: Coding & Agentic Frontier with MSA Architecture, 1M Context and Native Multimodality |
DataLearner blog | MiniMax M2.7 发布:模型开始帮自己训练自己 | Not provided |

MiniMax M3
MiniMaxAI