DataLearner logo
MI

MiMo-V2.6-Pro

Multimodal modelCoding model

MiMo-V2.6-Pro

Also known as: MiMo-V2.6-Pro-RL

Release date: 2026-09-22Views: 14
Live demoGitHubHugging FaceCompare
Parameters
1T
Context length
1M
Multilingual
Supported
Reasoning ability
5/5

MiMo-V2.6-Pro is Xiaomi's flagship omni-modal reasoning model, open-sourced under MIT in September 2026: a 1.02T-parameter sparse MoE with 42B active, 1M context, text/image/video/audio input, and 71.9 on DeepSWE v1.1.

Data sourced primarily from official releases (GitHub, Hugging Face, papers), then benchmark leaderboards, then third-party evaluators. Learn about our data methodology

MiMo-V2.6-Pro

Model basics

Reasoning traces
Supported
Thinking modes
Thinking modes not supported
Context length
1M tokens
Max output length
No data
Model type
Multimodal model
Modality (in / out)
Text, Image, Audio, Video → Text
Release date
2026-09-22
Model file size
No data
MoE architecture
Yes
Total params / Active params
1T / 42B
Knowledge cutoff
No data
MiMo-V2.6-Pro

Open source & experience

Code license
Weights license
MIT License- Commercial use permitted
GitHub repo
N/A
Live demo
N/A
MiMo-V2.6-Pro

Official resources

Paper
DataLearnerAI blog
N/A
MiMo-V2.6-Pro

API details

API speed
3/5
💡Default unit: $/1M tokens. If vendors use other units, follow their published pricing.
Standard
TypeConditionInputOutput
Text-¥3.00/ 1M tokens¥6.00/ 1M tokens
international
TypeConditionInputOutput
Text-$0.435/ 1M tokens$0.870/ 1M tokens
Cache PricingPrompt Cache
TypeTTLWriteRead
Text-$0.0036/ 1M tokens
Text-¥0.025/ 1M tokens

“—” means the modality is not billed in that direction, or the vendor has not published a price for it.

MiMo-V2.6-Pro

Benchmark Results

MiMo-V2.6-Pro currently shows benchmark results led by Terminal-Bench 2.1 (5 / 196, score 89.90), AutomationBench (2 / 20, score 53.10), DeepSWE (14 / 89, score 71.90). This page also consolidates core specs, context limits, and API pricing so you can evaluate the model from benchmark results and deployment constraints together.

Thinking

AI Agent - Tool Usage

5 evaluations
Benchmark / mode
Score
Rank/total
CyberGym
Thinking ModeTools
94
2 / 11
Terminal-Bench 2.1
Thinking ModeTools
89.90
5 / 196
OSWorld-Verified
Thinking ModeTools
82
6 / 28
Toolathlon-Verified
Thinking ModeTools
76.90
3 / 14
Terminal-Bench 4.0
Thinking ModeTools
34.90
18 / 91

Coding and Software Engineer

1 evaluations
Benchmark / mode
Score
Rank/total
DeepSWE
Thinking ModeTools
71.90
14 / 89

Agent Level Benchmark

2 evaluations
Benchmark / mode
Score
Rank/total
Job Bench
Thinking ModeTools
62
3 / 8
Agents' Last Exam
Thinking ModeTools
31.60
7 / 22

Productivity Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
AutomationBench
Thinking ModeTools
53.10
2 / 20
MiMo-V2.6-Pro

Publisher

MiMo-V2.6-Pro

Model Overview

MiMo-V2.6-Pro is the flagship omni-modal reasoning model that Xiaomi's MiMo team released and open-sourced in September 2026, the strongest member of the MiMo-V2.6 series. Xiaomi frames the series as a step toward scaling reinforcement learning into self-improvement: RL compute, environment and tool diversity, and grader compute are scaled together in one mixed run across coding, general agents, visual tasks and cybersecurity, rather than separate per-domain runs.


Architecture and specifications

The model is a sparse MoE with 1.02 trillion total parameters and 42B activated per token, 8 of 384 routed experts active; the language backbone has 70 layers (60 sliding-window attention + 10 global attention) with a hidden size of 6144. The models are natively omni-modal: they accept text, image, video and audio input and return text, with a 1M-token context window. The vision encoder is a 681M-parameter MiMo ViT (28 layers, 24 SWA + 4 full attention); the audio stack combines a 308M AudioTokenizer with a 127M audio patch encoder, and a 5-layer MTP speculative decoder predicts 7 tokens per pass.


Published benchmark results

Xiaomi reports the following for MiMo-V2.6-Pro: DeepSWE v1.1 71.9, Toolathlon-Verified 76.9, Terminal Bench 2.1 89.9, Terminal Bench 4.0 34.9, OSWorld-Verified 82.0, JobBench 62.0, Agents' Last Exam 31.6, AutomationBench v1.0.6 53.1, CyberGym 94.0, MiMo Code Bench 63.2, MiMo VisualCoding 72.3, GDPval-AA 2.1 1673. The official table also compares against Claude Opus 5, GPT-5.6 Sol and Claude Fable 5.


Pricing

On the Xiaomi MiMo platform, domestic pricing is ¥3 per million input tokens, ¥0.025 per million cached input tokens and ¥6 per million output tokens; international pricing is $0.435 input, $0.0036 cached input and $0.87 output per million tokens, with a 50% discount for batch inference.


Release and license

The MiMo-V2.6 series was published on Hugging Face on September 21, 2026 under the MIT license, and served through the Xiaomi MiMo platform (mimo.mi.com) from September 22, with distribution also via ModelScope and OpenRouter. Xiaomi published a live dashboard of the RL post-training run: Pro and Flash each took under six days and 30 steps, about 750,000 trajectories in total, at a reported cost of roughly $2.62M and $0.85M respectively.

MiMo-V2.6-Pro

FAQ

Is MiMo-V2.6-Pro open source?

Yes. Xiaomi published the weights on Hugging Face as XiaomiMiMo/MiMo-V2.6-Pro-RL on September 21, 2026 under the MIT license.

How large is MiMo-V2.6-Pro?

It is a sparse MoE with 1.02 trillion total parameters and 42B activated per token, 8 of 384 routed experts active.

What modalities and context length does it support?

It natively accepts text, image, video and audio input, returns text, and supports a 1M-token context window.

How much does the MiMo-V2.6-Pro API cost?

International pricing is $0.435 per million input tokens, $0.0036 per million cached input tokens and $0.87 per million output tokens; batch inference is 50% off.

DataLearner on WeChat

Follow DataLearner on WeChat for AI model updates and research notes.

DataLearner WeChat QR code