Xiaomi Launches MiMo-V2.5 With 1M Context Window at $0.40 per Million Input Tokens
Xiaomi released MiMo-V2.5 on April 22, 2026, a native omnimodal model with a 1,048,576 token context window. The model is priced at $0.40 per million input tokens and $2 per million output tokens, positioning it as a cost-efficient alternative for agentic applications requiring multimodal perception across image and video understanding.
MiMo-V2.5 — Quick Specs
Xiaomi Launches MiMo-V2.5 With 1M Context Window at $0.40 per Million Input Tokens
Xiaomi released MiMo-V2.5 on April 22, 2026, a native omnimodal model featuring a 1,048,576 token context window priced at $0.40 per million input tokens and $2 per million output tokens.
Specifications and Pricing
MiMo-V2.5 offers:
- Context window: 1,048,576 tokens (1M)
- Input pricing: $0.40 per million tokens
- Output pricing: $2 per million tokens
- Release date: April 22, 2026
According to Xiaomi, the model delivers "Pro-level agentic performance at roughly half the inference cost" compared to unspecified alternatives, though the company has not provided independent benchmark scores to verify these claims.
Technical Capabilities
Xiaomi describes MiMo-V2.5 as a "native omnimodal model" designed for multimodal perception across image and video understanding tasks. The company claims the model surpasses its predecessor, MiMo-V2-Omni, in multimodal perception, though specific benchmark comparisons were not disclosed.
The 1M context window is designed to handle complete documents, extended conversations, and complex task contexts in a single inference pass. Xiaomi positions this capability as particularly suited for integration with agent frameworks.
Availability
The model is currently available through OpenRouter, which routes requests across multiple providers to optimize uptime and handle varying prompt sizes. OpenRouter supports the model's reasoning capabilities through a dedicated reasoning parameter that exposes step-by-step thinking processes via a reasoning_details array in API responses.
What This Means
MiMo-V2.5 enters an increasingly competitive omnimodal model market with a clear value proposition: extended context at lower input pricing than many enterprise-tier alternatives. At $0.40 per million input tokens, it undercuts several comparable models while offering a 1M context window—a specification typically reserved for premium tiers.
The focus on agentic workflows suggests Xiaomi is targeting developers building autonomous systems that require sustained reasoning across multimodal inputs. However, without published benchmark scores on standard evaluation sets like MMLU, VQAv2, or video understanding benchmarks, independent assessment of the model's claimed performance advantages remains difficult. The model's effectiveness will ultimately be determined by real-world deployment results in production agent systems.
Related Articles
Meta Releases Muse Spark 1.3 Contributor, a Low-Cost Multimodal Reasoning Model With 1M Context Window
Meta has released Muse Spark 1.3 Contributor, described as the cost-efficient contributor tier of its multimodal reasoning model line. The model offers a 1 million token context window at $0.10 per 1M input tokens and $0.20 per 1M output tokens, targeting experimentation and early-stage agentic workflows.
Meta Releases Muse Spark 1.3, a Free Multimodal Reasoning Model with 1M-Token Context
Meta has released Muse Spark 1.3, a multimodal reasoning model with a 1M-token context window, listed as free on OpenRouter. The model targets long-running agentic, multi-agent, and coding workflows, though audio input support remains incomplete.
Alibaba Releases Qwen3.8 Max (0902), a 2.4-Trillion-Parameter MoE Model With 1M-Token Context
Alibaba's Qwen team released Qwen3.8 Max (0902), a 2.4-trillion-parameter mixture-of-experts model with a 1M-token context window that accepts text, image, and video input. The snapshot is post-trained for coding, agentic workflows, and long-horizon task execution, priced at $2/$6 per 1M input/output tokens.
InclusionAI Releases Ling 3.0 Flash Fin, a Finance-Focused MoE Model with 5.1B Active Parameters
InclusionAI has released Ling 3.0 Flash Fin, a finance-specialized mixture-of-experts model built on Ling 3.0 Flash. The model activates 5.1B of its 124B total parameters and targets long-horizon investment planning tasks while retaining general reasoning, coding, and math capabilities.
Comments
Loading...