DataLearner logo

MiniMax M3 Benchmark Details

MiniMax M3 currently shows benchmark results led by AA-LCR (1 / 27, score 80.33), BrowseComp (12 / 54, score 83.50), SWE-Bench Pro - Public (15 / 59, score 59). This page also compares it with 5 competitor models and 4 predecessor or same-series models, including performance and pricing views when available. 1 source link is attached for reference.

Benchmark Results

MiniMax M3

Benchmark Results

Thinking
Tool usage
Internet

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
81.31
132 / 270

Coding and Software Engineer

4 evaluations
Benchmark / mode
Score
Rank/total
Text Arena (Coding)
Standard Mode
1527.75
14 / 35
SWE-Bench Pro - Public
Thinking ModeTools
59
15 / 59
SciCode
Thinking Mode
45.37
10 / 14
PostTrain Bench
Thinking ModeTools
37
2 / 5

AI Agent - Information Search

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
Thinking ModeToolsInternet
83.50
12 / 54

General Knowledge

2 evaluations
Benchmark / mode
Score
Rank/total
LiveBench
Deep Thinking Mode
70.02
40 / 115
45.40
8 / 12

Text Embedding

1 evaluations
Benchmark / mode
Score
Rank/total
Context Arena
Thinking Mode
51.15
87 / 126

Long Context

1 evaluations
Benchmark / mode
Score
Rank/total
AA-LCR
Thinking Mode
80.33
1 / 27

AI Agent - Tool Usage

3 evaluations
Benchmark / mode
Score
Rank/total
MCP-Atlas
Thinking ModeTools
74.20
24 / 40
OSWorld-Verified
Thinking ModeTools
70
19 / 26
Terminal-Bench 2.1
Thinking ModeTools
66
38 / 47

Productivity Knowledge

3 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA v2
Thinking ModeTools
1379.75
16 / 22
AA-Briefcase
Thinking ModeTools
1107.16
14 / 19
AA-AnalystAgent
Thinking ModeToolsInternet
10
11 / 12

Agent Level Benchmark

1 evaluations
Benchmark / mode
Score
Rank/total
τ³-Banking
Thinking ModeTools
15.26
10 / 12

Competitor Comparison

Benchmark scores for MiniMax M3 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

11 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkMiniMax M3CurrentGLM-5.2Kimi K2.6Qwen3.7 MaxGLM 5.1
GPQA Diamond
科学与综合推理
81.31Standard Mode
91.86Thinking Level · High
90.50Thinking Enabled
92.40Thinking Level · High
86.20Thinking Enabled
PostTrain Bench
编程与软件工程
37.00Thinking Enabled | Tools
34.30Thinking Level · High | Tools
--
--
--
SciCode
编程与软件工程
45.37Thinking Enabled
--
--
53.50Thinking Level · High
--
SWE-Bench Pro - Public
编程与软件工程
59.00Thinking Enabled | Tools
62.10Thinking Enabled | Tools
58.60Thinking Enabled | Tools
60.60Thinking Enabled | Tools
58.40Thinking Enabled | Tools
Text Arena (Coding)
编程与软件工程
1527.75Standard Mode
1593.25Thinking Level · High
--
1540.77Standard Mode
1534.00Standard Mode
BrowseComp
AI Agent - 信息收集
83.50Thinking Enabled | Tools
--
83.20Thinking Enabled | Tools
--
79.30Thinking Enabled | Tools
LiveBench
综合评估
70.02Deep Thinking Mode
76.24Standard Mode
72.17Thinking Enabled
74.29Deep Thinking Mode
70.18Standard Mode
Context Arena
文本向量检索
51.15Thinking Enabled
72.34Thinking Level · High
64.63Thinking Enabled
56.01Standard Mode
62.05Thinking Enabled
MCP-Atlas
AI Agent - 工具使用
74.20Thinking Enabled | Tools
76.80Thinking Enabled | Tools
69.40Thinking Enabled | Tools
76.40Thinking Enabled | Tools
75.60Standard Mode | Tools
OSWorld-Verified
AI Agent - 工具使用
70.00Thinking Enabled | Tools
--
73.10Thinking Enabled | Tools
--
--
Terminal-Bench 2.1
AI Agent - 工具使用
66.00Thinking Enabled | Tools
81.00Thinking Level · High | Tools
53.56Thinking Enabled
--
58.70Thinking Level · High | Tools

Standard API Pricing: MiniMax M3 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

MiniMax M3
Supplier: MiniMaxAI
Standard input: ¥2.1 / 1M tokens
Standard output: ¥8.4 / 1M tokens
GLM-5.2
Supplier: 智谱AI
Standard input: $1.4 / 1M tokens
Standard output: $4.4 / 1M tokens
Kimi K2.6
Supplier: Facebook AI研究实验室
Standard input: $0.95 / 1M tokens
Standard output: $4 / 1M tokens
Qwen3.7 Max
Supplier: 阿里巴巴
Standard input: ¥12 / 1M tokens
Standard output: ¥36 / 1M tokens
GLM 5.1
Supplier: 智谱AI
Standard input: $1.4 / 1M tokens
Standard output: $4.4 / 1M tokens
Qwen3.7-Max-Preview
Supplier: 阿里巴巴
Standard input: $2.5 / 1M tokens
Standard output: $7.5 / 1M tokens
ModelSupplierStandard inputStandard outputBase price applies to
MiniMax M3
MiniMaxAI¥2.1 / 1M tokens¥8.4 / 1M tokens
GLM-5.2
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens
Kimi K2.6
Facebook AI研究实验室$0.95 / 1M tokens$4 / 1M tokens
Qwen3.7 Max
阿里巴巴¥12 / 1M tokens¥36 / 1M tokens
GLM 5.1
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens
Qwen3.7-Max-Preview
阿里巴巴$2.5 / 1M tokens$7.5 / 1M tokens

Version History

How each version of the MiniMax M3 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

7 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkMiniMax M3CurrentMiniMax-M2.7MiniMax M2.5M2.1
GPQA Diamond
科学与综合推理
81.31Standard Mode
87.00Thinking Enabled
85.20Thinking Enabled
81.00Thinking Enabled
SWE-Bench Pro - Public
编程与软件工程
59.00Thinking Enabled | Tools
56.20Thinking Enabled | Tools
55.40Thinking Enabled | Tools
32.60Thinking Enabled | Tools
BrowseComp
AI Agent - 信息收集
83.50Thinking Enabled | Tools
--
76.30Thinking Enabled | Tools
47.40Thinking Enabled | Tools
LiveBench
综合评估
70.02Deep Thinking Mode
63.49Deep Thinking Mode
60.14Deep Thinking Mode
--
Context Arena
文本向量检索
51.15Thinking Enabled
33.29Thinking Enabled
--
--
AA-LCR
长上下文能力
80.33Thinking Enabled
69.00Thinking Enabled | Tools
69.50Thinking Enabled
--
GDPval-AA v2
生产力知识
1379.75Thinking Enabled | Tools
1495.00Thinking Enabled | Tools
--
--

Single-Benchmark Version Trend

Viewing: GPQA Diamond · 科学与综合推理

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the MiniMax M3 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

MiniMax M3
Supplier: MiniMaxAI
Standard input: ¥2.1 / 1M tokens
Standard output: ¥8.4 / 1M tokens
MiniMax-M2.7
Supplier: MiniMaxAI
Standard input: $0.3 / 1M tokens
Standard output: $1.2 / 1M tokens
MiniMax M2.5
Supplier: MiniMaxAI
Standard input: $0.3 / 1M tokens
Standard output: $2.4 / 1M tokens
M2.1
Supplier: MiniMaxAI
Standard input: ¥2.1 / 1M tokens
Standard output: ¥8.4 / 1M tokens
ModelSupplierStandard inputStandard outputBase price applies to
MiniMax M3
MiniMaxAI¥2.1 / 1M tokens¥8.4 / 1M tokens
MiniMax-M2.7
MiniMaxAI$0.3 / 1M tokens$1.2 / 1M tokens
MiniMax M2.5
MiniMaxAI$0.3 / 1M tokens$2.4 / 1M tokens
M2.1
MiniMaxAI¥2.1 / 1M tokens¥8.4 / 1M tokens

Sources