Xiaomi Launches MiMo-V2.6-Pro-UltraSpeed: Same Quality, 10x Faster Output
Xiaomi's MiMo-V2.6-Pro-UltraSpeed is a fast-inference edition of the company's 1T-parameter flagship MiMo-V2.6-Pro, delivering roughly 10x the output speed at matching quality. It retains the 1M-token context window and native multimodal capabilities, priced at $4.35/$8.70 per 1M input/output tokens.
MiMo-V2.6-Pro-UltraSpeed — Quick Specs
Xiaomi released MiMo-V2.6-Pro-UltraSpeed on September 21, 2026, a fast-inference variant of its flagship MiMo-V2.6-Pro model. According to Xiaomi, the model is built from the same 1-trillion-parameter checkpoint as MiMo-V2.6-Pro and matches it in output quality while delivering roughly 10x the throughput.
The model carries a 1-million-token context window and native multimodal capabilities, inherited unchanged from its parent model. Xiaomi positions it for agentic workflows, claiming top-tier performance across coding, visual, general, and research tasks, with particular emphasis on complex, long-horizon tasks and compatibility across a range of agent harnesses.
Pricing and performance
MiMo-V2.6-Pro-UltraSpeed is priced at $4.35 per 1M input tokens and $8.70 per 1M output tokens — exactly 10x the pricing of the base MiMo-V2.6-Pro model, which runs at $0.435/$0.87 per 1M tokens. Per OpenRouter data, the model shows a P50 latency of 2.32 seconds and throughput of 83 tokens per second from Xiaomi's own endpoint, with 92.30% uptime over the last 24 hours across providers.
Early usage on OpenRouter shows adoption through tools including Open WebUI, Cherry Studio, Cline, and an agent framework called Dakota Agent, though volumes remain modest (in the range of hundreds to a few thousand tokens processed per integration in the observed window).
Context within Xiaomi's MiMo lineup
MiMo-V2.6-Pro-UltraSpeed is the latest addition to Xiaomi's expanding MiMo model family, which now includes MiMo-V2.6-Flash (309B total parameters, 15B active, Mixture-of-Experts, priced at $0.14/$0.28 per 1M tokens), MiMo-V2.6-Pro, MiMo-V2.5-Pro, MiMo-V2.5, MiMo-V2-Omni, MiMo-V2-Pro, and MiMo-V2-Flash. Xiaomi claims MiMo-V2-Flash ranks as the top open-source model globally on SWE-bench Verified and SWE-bench Multilingual, with performance comparable to Claude Sonnet 4.5 at about 3.5% of the cost — a claim not independently verified.
No independent benchmark scores for MiMo-V2.6-Pro-UltraSpeed itself have been published. Xiaomi has not disclosed a training data cutoff date for the underlying MiMo-V2.6-Pro checkpoint.
What this means
This release is not a new model in the traditional sense — it's the same MiMo-V2.6-Pro weights served through an inference stack optimized for speed, at a 10x price premium to match. That trade-off signals Xiaomi is targeting latency-sensitive agentic applications (multi-step tool use, real-time coding assistants) where throughput matters more than per-token cost. The pricing structure — linear 10x scaling on both cost and speed — suggests this is a distinct serving tier rather than a distilled or quantized variant, though Xiaomi has not detailed the underlying inference optimization (e.g., speculative decoding, custom hardware, or parallel decoding). With 92.30% uptime reported across providers in its first day of availability, the model is still in an early rollout phase, and real-world throughput and reliability at scale remain to be seen. For teams building latency-critical agents, this offers an alternative to premium Western models like Claude or GPT at potentially competitive pricing, but the lack of independent benchmarks means quality claims should be treated cautiously until third-party evaluations emerge.
Related Articles
Xiaomi Launches MiMo-V2.6-Pro, a 1T+ Parameter Model with 1M-Token Context
Xiaomi has released MiMo-V2.6-Pro, a flagship foundation model exceeding 1 trillion parameters with a 1M-token context window and native multimodal support. The model is priced at $0.435 per 1M input tokens and $0.87 per 1M output tokens, targeting agentic and long-horizon tasks.
Xiaomi Releases MiMo-V2.6-Flash: Open-Source MoE Model with 1M-Token Context, $0.14/$0.28 per 1M Tokens
Xiaomi has released MiMo-V2.6-Flash, an open-source Mixture-of-Experts model with 309B total parameters and 15B activated per token, featuring a 1M-token context window and native multimodal capabilities. Priced at $0.14 per 1M input tokens and $0.28 per 1M output tokens, it targets agentic coding and long-horizon task workflows.
xAI Ships Grok 4.7, Cuts Price to $1.60/$4.80 per 1M Tokens With 500K Context
xAI has released Grok 4.7, the successor to Grok 4.6, listed on OpenRouter with a 500K token context window and pricing of $1.60 per 1M input tokens and $4.80 per 1M output tokens. The company claims improvements in long-running software engineering tasks, self-verification, and professional document drafting.
Z.ai Releases GLM-5.3-FlashX, a 200 Tokens/Second Variant of Its GLM-5.3-Flash Model
Z.ai has released GLM-5.3-FlashX, a high-speed variant of GLM-5.3-Flash built on a hybrid sparse and linear attention architecture with 320B total parameters (18B active). The model supports a 1M-token context window and claims inference speeds of up to 200 tokens per second.
Comments
Loading...