model releaseXiaomi

Xiaomi Launches MiMo-V2.6-Pro, a 1T+ Parameter Model with 1M-Token Context

TL;DR

Xiaomi has released MiMo-V2.6-Pro, a flagship foundation model exceeding 1 trillion parameters with a 1M-token context window and native multimodal support. The model is priced at $0.435 per 1M input tokens and $0.87 per 1M output tokens, targeting agentic and long-horizon tasks.

2 min read
0

MiMo-V2.6-Pro — Quick Specs

Context window1000K tokens
Input$0.435/1M tokens
Output$0.87/1M tokens

Xiaomi has released MiMo-V2.6-Pro, the newest version of its flagship foundation model, built at a scale of over 1 trillion parameters. The model is available now via OpenRouter and is listed with a release date of September 21, 2026.

Specifications

MiMo-V2.6-Pro features a 1-million-token context window and native multimodal input handling, according to Xiaomi. Pricing on OpenRouter is set at $0.435 per 1M input tokens and $0.87 per 1M output tokens, with cache-read pricing listed at $0.0036 per 1M tokens.

OpenRouter's live telemetry shows a P50 latency of 1.40 seconds and throughput of 88 tokens per second, with 99.48% uptime recorded over the trailing 24-hour window and 100% availability across all locations over three days.

Xiaomi describes the model as optimized for agentic workflows, claiming top-tier performance across coding, visual, general, and research scenarios. The company states MiMo-V2.6-Pro excels at complex, long-horizon tasks with robust generalization across a range of agent harnesses — no independent benchmark scores for this specific version have been published.

Part of a broader model family

The release accompanies several sibling models in Xiaomi's MiMo-V2.6 line:

  • MiMo-V2.6-Pro-UltraSpeed — built from the same 1T-parameter checkpoint, claimed by Xiaomi to match Pro-level quality at roughly 10x the output speed, priced at $4.35/$8.70 per 1M tokens.
  • MiMo-V2.6-Flash — a Mixture-of-Experts model with 309B total parameters and 15B active parameters per token, using hybrid attention, priced at $0.14/$0.28 per 1M tokens.

These sit alongside earlier releases including MiMo-V2.5-Pro (1.1M context, $0.3045/$0.609 per 1M), which Xiaomi says has topped benchmarks including ClawEval, GDPVal, and SWE-bench Pro, and MiMo-V2-Flash, which the company claims ranks first among open-source models on SWE-bench Verified and SWE-bench Multilingual while costing roughly 3.5% as much as Claude Sonnet 4.5.

What this means

Xiaomi is positioning MiMo-V2.6-Pro as a low-cost alternative to Western frontier models for agentic and long-context workloads, undercutting typical flagship pricing by a wide margin while claiming comparable capability. The trillion-plus parameter count and 1M-token context are notable on paper, but Xiaomi has not published independent benchmark scores for this specific version — buyers evaluating it for production agent pipelines should treat the company's capability claims as unverified until third-party evaluations appear. The accompanying UltraSpeed and Flash variants suggest Xiaomi is building out a tiered product line similar to competitors, trading cost and latency against raw model size for different workload profiles.

Related Articles

model release

Xiaomi Launches MiMo-V2.6-Pro-UltraSpeed: Same Quality, 10x Faster Output

Xiaomi's MiMo-V2.6-Pro-UltraSpeed is a fast-inference edition of the company's 1T-parameter flagship MiMo-V2.6-Pro, delivering roughly 10x the output speed at matching quality. It retains the 1M-token context window and native multimodal capabilities, priced at $4.35/$8.70 per 1M input/output tokens.

model release

xAI Ships Grok 4.7, Cuts Price to $1.60/$4.80 per 1M Tokens With 500K Context

xAI has released Grok 4.7, the successor to Grok 4.6, listed on OpenRouter with a 500K token context window and pricing of $1.60 per 1M input tokens and $4.80 per 1M output tokens. The company claims improvements in long-running software engineering tasks, self-verification, and professional document drafting.

model release

Xiaomi Releases MiMo-V2.6-Flash: Open-Source MoE Model with 1M-Token Context, $0.14/$0.28 per 1M Tokens

Xiaomi has released MiMo-V2.6-Flash, an open-source Mixture-of-Experts model with 309B total parameters and 15B activated per token, featuring a 1M-token context window and native multimodal capabilities. Priced at $0.14 per 1M input tokens and $0.28 per 1M output tokens, it targets agentic coding and long-horizon task workflows.

model release

Z.ai Releases GLM-5.3-FlashX, a 200 Tokens/Second Variant of Its GLM-5.3-Flash Model

Z.ai has released GLM-5.3-FlashX, a high-speed variant of GLM-5.3-Flash built on a hybrid sparse and linear attention architecture with 320B total parameters (18B active). The model supports a 1M-token context window and claims inference speeds of up to 200 tokens per second.

Comments

Loading...