Xiaomi's MiMo-V2.6-Pro Becomes Top Open-Weights Model, Trained for $3M According to Xiaomi
Xiaomi released MiMo-V2.6-Pro, a 1.02T-parameter mixture-of-experts model with 42B active parameters, which debuted as the top-scoring open-weights model on Artificial Analysis' Intelligence Index (46). The company claims the model's RL training run cost roughly $2.6M and completed in 130 hours.
Xiaomi Enters the Frontier Open-Weights Race
Xiaomi released MiMo-V2.6-Pro on September 21, 2026, a mixture-of-experts model with 1.02 trillion total parameters and 42 billion active parameters per forward pass. According to Artificial Analysis, the model debuted as the top-scoring open-weights model on its Intelligence Index, posting a score of 46 — the highest recorded for an openly-licensed model to date.
The release includes two additional variants: MiMo-V2.6-Flash, described by Xiaomi as balancing intelligence, efficiency, and cost, and MiMo-V2.6-Pro-UltraSpeed, which the company claims delivers up to 20x faster output generation at equivalent quality. All models are released under the MIT license, according to Hugging Face contributor Victor Mustar.
Pricing and Cost Efficiency
Artificial Analysis lists API pricing for MiMo-V2.6-Pro at $0.435 per 1M input tokens and $0.87 per 1M output tokens. On a cost-per-task basis, Artificial Analysis reports the model lands on its Intelligence-vs-Cost Pareto frontier at $0.13 per Intelligence Index task — meaning no other model currently available delivers comparable intelligence at a lower cost.
Training Cost Claims
Xiaomi's technical report and follow-on commentary from researchers put the total training cost around $3M, with the final reinforcement learning run specifically costing an estimated $2.6M, according to a figure cited by AI researcher Zephyr (@zephyr_z9) on X. That RL run reportedly ran for 130 hours and processed 75 billion tokens. The team says it scaled RL compute along three axes: larger batch sizes (1,568 samples per update) with training contexts up to 1 million tokens and 3.5–3.7 billion tokens per step; a broader multi-task training suite spanning coding, agents, vision, cybersecurity, and music; and increased grader compute for more precise reward signals on long-horizon tasks.
The RL infrastructure reportedly runs on JAX and TPUs, with researcher Tianjun Zhang noting that scaling the system is largely a configuration change rather than a code rewrite — a claim that, if accurate, suggests Xiaomi's RL pipeline is unusually portable and cheap to scale further.
Native Omnimodality and Open Tooling
Xiaomi describes MiMo-V2.6-Pro and Flash as "natively omnimodal," trained jointly across text, vision, code, cybersecurity vulnerability reproduction, and music generation tasks rather than bolting modalities onto a text-only base. The company says gains in one domain reinforce performance in others due to shared RL environments.
Xiaomi is open-sourcing the environment code and training recipes behind the model, covering coding/software engineering, cyber vulnerability reproduction (built on the ARVO environment), general knowledge work, web development, and music scoring. The full training datasets — reportedly more than 7,000 RL tasks — have not yet been released.
Background
The release follows a period of unusual transparency from Xiaomi: Fuli Luo, a former DeepSeek engineer now at Xiaomi, publicly livestreamed the model's final RL training runs in the days prior to launch, an atypical move for a frontier lab. Xiaomi is not traditionally grouped among China's established "AI Tiger" labs, making its emergence at the top of an open-weights leaderboard notable to close observers of the Chinese AI ecosystem.
What This Means
MiMo-V2.6-Pro's headline claim — frontier-adjacent intelligence at a fraction of typical training and inference cost — is unverified beyond Xiaomi's own report and third-party estimates circulating on social media. If the $2.6M RL run figure holds up under independent scrutiny, it reinforces a broader trend: post-training and RL infrastructure, not just pretraining scale, is becoming the primary lever for closing the gap between open and closed frontier models. Xiaomi's decision to open-source its RL environments rather than just weights could matter more long-term than the model itself, since reusable, high-quality RL environments may prove as valuable to future labs as pretraining corpora were in the previous cycle.
Related Articles
Xiaomi Releases MiMo-V2.6-Pro-RL, a 1.02T-Parameter Omnimodal Model with 1M-Token Context
Xiaomi's MiMo team has released MiMo-V2.6-Pro-RL, a 1.02-trillion-parameter sparse mixture-of-experts model with 42B active parameters, 1M-token context, and native text/image/video/audio processing. The model was trained via a single mixed reinforcement learning run spanning coding, agentic, visual, and cybersecurity tasks, with benchmark scores that Xiaomi claims approach or match Claude Opus 5 and GPT-5.6 on several agentic and coding tests.
Xiaomi Releases MiMo-V2.6-Flash-RL, a 309B-Parameter MoE Model with 1M-Token Context and Native Omnimodal Support
Xiaomi's MiMo team released MiMo-V2.6-Flash-RL, an efficiency-tier checkpoint in the MiMo-V2.6 series featuring a 309B-parameter (15B active) Mixture-of-Experts architecture, 1M-token context, and native support for text, image, video, and audio. The model uses a single mixed reinforcement learning run across coding, agentic, visual, and cybersecurity tasks rather than domain-specific training.
Xiaomi Launches MiMo-V2.6-Pro-UltraSpeed: Same Quality, 10x Faster Output
Xiaomi's MiMo-V2.6-Pro-UltraSpeed is a fast-inference edition of the company's 1T-parameter flagship MiMo-V2.6-Pro, delivering roughly 10x the output speed at matching quality. It retains the 1M-token context window and native multimodal capabilities, priced at $4.35/$8.70 per 1M input/output tokens.
Xiaomi Releases MiMo-V2.6-Flash: Open-Source MoE Model with 1M-Token Context, $0.14/$0.28 per 1M Tokens
Xiaomi has released MiMo-V2.6-Flash, an open-source Mixture-of-Experts model with 309B total parameters and 15B activated per token, featuring a 1M-token context window and native multimodal capabilities. Priced at $0.14 per 1M input tokens and $0.28 per 1M output tokens, it targets agentic coding and long-horizon task workflows.
Comments
Loading...