Is MiMo-V2.6-Flash open source?
Yes. The weights are on Hugging Face as XiaomiMiMo/MiMo-V2.6-Flash-RL under the MIT license, published September 21, 2026.
MiMo-V2.6-Flash
Also known as: MiMo-V2.6-Flash-RL
MiMo-V2.6-Flash is Xiaomi's efficient omni-modal reasoning model, open-sourced under MIT in September 2026: a 309B sparse MoE with 15B active, 1M context, text/image/video/audio input, and 67.9 on DeepSWE v1.1.
Data sourced primarily from official releases (GitHub, Hugging Face, papers), then benchmark leaderboards, then third-party evaluators. Learn about our data methodology
| Type | Condition | Input | Output |
|---|---|---|---|
| Text | - | ¥1.00/ 1M tokens | ¥2.00/ 1M tokens |
| Type | Condition | Input | Output |
|---|---|---|---|
| Text | - | $0.140/ 1M tokens | $0.280/ 1M tokens |
| Type | TTL | Write | Read |
|---|---|---|---|
| Text | - | — | $0.0028/ 1M tokens |
| Text | - | — | ¥0.020/ 1M tokens |
“—” means the modality is not billed in that direction, or the vendor has not published a price for it.
MiMo-V2.6-Flash currently shows benchmark results led by CyberGym (1 / 11, score 95.10), Terminal-Bench 2.1 (18 / 196, score 87.60), AutomationBench (3 / 20, score 52.30). This page also consolidates core specs, context limits, and API pricing so you can evaluate the model from benchmark results and deployment constraints together.
MiMo-V2.6-Flash is the efficient omni-modal reasoning model Xiaomi's MiMo team released and open-sourced in September 2026 alongside the flagship MiMo-V2.6-Pro. It targets high-frequency and high-volume workloads at low cost and uses the same mixed reinforcement-learning recipe.
The model is a sparse MoE with 309B total parameters and 15B activated per token, 8 of 256 routed experts active; the language backbone has 48 layers (39 sliding-window attention + 9 global attention) with a hidden size of 4096. The models are natively omni-modal: they accept text, image, video and audio input and return text, with a 1M-token context window. The vision encoder is a 681M-parameter MiMo ViT (28 layers, 24 SWA + 4 full attention); the audio stack combines a 308M AudioTokenizer with a 127M audio patch encoder, and a 5-layer MTP speculative decoder predicts 7 tokens per pass.
Xiaomi reports the following for MiMo-V2.6-Flash: DeepSWE v1.1 67.9, Toolathlon-Verified 73.6, Terminal Bench 2.1 87.6, Terminal Bench 4.0 28.8, OSWorld-Verified 80.8, JobBench 61.2, Agents' Last Exam 27.6, AutomationBench v1.0.6 52.3, CyberGym 95.1, MiMo Code Bench 61.2, MiMo VisualCoding 71.5. The official table also compares against Claude Opus 5, GPT-5.6 Sol and Claude Fable 5.
On CyberGym, Flash scores 95.1, slightly above Pro's 94.0.
Domestic pricing is ¥1 per million input tokens, ¥0.02 per million cached input tokens and ¥2 per million output tokens; international pricing is $0.14 input, $0.0028 cached input and $0.28 output per million tokens, with a 50% discount for batch inference.
The MiMo-V2.6 series was published on Hugging Face on September 21, 2026 under the MIT license, and served through the Xiaomi MiMo platform (mimo.mi.com) from September 22, with distribution also via ModelScope and OpenRouter. Xiaomi published a live dashboard of the RL post-training run: Pro and Flash each took under six days and 30 steps, about 750,000 trajectories in total, at a reported cost of roughly $2.62M and $0.85M respectively.
Yes. The weights are on Hugging Face as XiaomiMiMo/MiMo-V2.6-Flash-RL under the MIT license, published September 21, 2026.
Flash is a 309B sparse MoE with 15B active parameters versus Pro's 1.02T/42B. It trails Pro on most agentic benchmarks (67.9 vs 71.9 on DeepSWE v1.1) but scores slightly higher on CyberGym (95.1 vs 94.0) at roughly a third of the price.
International pricing is $0.14 per million input tokens, $0.0028 per million cached input tokens and $0.28 per million output tokens.
Follow DataLearner on WeChat for AI model updates and research notes.
