DataLearner logo

MiniMax M3vsStep 3.7 Flash

Across 3 shared benchmarks, MiniMax M3 leads overall: MiniMax M3 wins 3, Step 3.7 Flash wins 0, with 0 ties and an average score difference of +5.63.

MiniMaxAI
MiniMax M3

MiniMaxAI · 2026-06-01 · Multimodal model

StepFunAI
Step 3.7 Flash

StepFunAI · 2026-05-29 · Reasoning model

MiniMax M33 wins(100%)(0%)0 winsStep 3.7 Flash

Benchmark scores

Grouped by capability, sorted by largest gap within each. 3 shared benchmarks.

AI Agent - Information Search

MiniMax M3 1/1
BenchmarkMiniMax M3Step 3.7 FlashDiff
BrowseComp83.5012 / 54Thinking (With Tools + Internet)75.8225 / 54Thinking (With Tools)+7.68

AI Agent - Tool Usage

MiniMax M3 1/1
BenchmarkMiniMax M3Step 3.7 FlashDiff
Terminal-Bench 2.16635 / 44Thinking (With Tools)59.5037 / 44Thinking (With Tools)+6.50

Coding and Software Engineer

MiniMax M3 1/1
BenchmarkMiniMax M3Step 3.7 FlashDiff
SWE-Bench Pro - Public5913 / 57Thinking (With Tools)56.3025 / 57Thinking (With Tools)+2.70

Specs

FieldMiniMax M3Step 3.7 Flash
PublisherMiniMaxAIStepFunAI
Release date2026-06-012026-05-29
Model typeMultimodal modelReasoning model
ArchitectureMoEMoE
Parameters428B198B
Context length1M256K
Max output512KNot available

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemMiniMax M3Step 3.7 Flash
Text input¥2.1 / 1M tokens¥1.35 / 1M tokens
Text output¥8.4 / 1M tokens¥8.1 / 1M tokens
Cache read¥0.42 / 1M tokens¥0.27 / 1M tokens

Summary

  • MiniMax M3leads in:AI Agent - Information Search (1/1), AI Agent - Tool Usage (1/1), Coding and Software Engineer (1/1)

On average across the 3 shared benchmarks, MiniMax M3 scores 5.63 higher.

Largest single-benchmark gap: BrowseComp — MiniMax M3 83.50 vs Step 3.7 Flash 75.82 (+7.68).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.