model releaseXiaomi

Xiaomi Releases MiMo-V2.6-Flash: Open-Source MoE Model with 1M-Token Context, $0.14/$0.28 per 1M Tokens

TL;DR

Xiaomi has released MiMo-V2.6-Flash, an open-source Mixture-of-Experts model with 309B total parameters and 15B activated per token, featuring a 1M-token context window and native multimodal capabilities. Priced at $0.14 per 1M input tokens and $0.28 per 1M output tokens, it targets agentic coding and long-horizon task workflows.

2 min read
0

MiMo-V2.6-Flash — Quick Specs

Context window1000K tokens
Input$0.14/1M tokens
Output$0.28/1M tokens

Xiaomi has released MiMo-V2.6-Flash, an open-source foundation model built on a Mixture-of-Experts (MoE) architecture with 309 billion total parameters, of which 15 billion are activated per token. The model is available now via OpenRouter and other providers, priced at $0.14 per 1 million input tokens and $0.28 per 1 million output tokens.

Specifications

MiMo-V2.6-Flash features a 1-million-token context window and native multimodal capabilities, handling text alongside visual inputs within a unified architecture. According to Xiaomi, the model employs a hybrid attention mechanism designed to improve computational efficiency relative to standard transformer attention at this parameter scale.

The company positions the model for agentic workflows, claiming strong performance across coding, visual, general, and research scenarios. Xiaomi states the model excels at "complex, long-horizon tasks with robust generalization across a diverse range of agent harnesses" — though no specific benchmark scores (e.g., SWE-bench, MMLU) were disclosed for this particular release, unlike its predecessor MiMo-V2-Flash, which Xiaomi says ranked #1 among open-source models on SWE-bench Verified and SWE-bench Multilingual.

On OpenRouter, the model shows a P50 latency of 1.04 seconds and throughput of 116 tokens per second from Xiaomi's own hosting, with 99.81% uptime recorded over the platform's tracking window.

Where it fits in Xiaomi's lineup

MiMo-V2.6-Flash is the latest entry in Xiaomi's expanding MiMo model family, which now spans multiple tiers:

  • MiMo-V2.6-Pro: flagship model exceeding 1 trillion parameters, $0.435/$0.87 per 1M tokens
  • MiMo-V2.6-Pro-UltraSpeed: same 1T-parameter checkpoint as Pro but claims ~10x output speed, $4.35/$8.70 per 1M tokens
  • MiMo-V2.5-Pro: prior flagship, $0.3045/$0.609 per 1M tokens
  • MiMo-V2.5: omnimodal model, $0.119/$0.238 per 1M tokens
  • MiMo-V2-Flash: predecessor MoE model with the same 309B/15B parameter split but a 256K context window

The V2.6-Flash release keeps the same MoE parameter configuration as V2-Flash while expanding context from 256K to 1M tokens and adding native multimodal input handling.

What this means

Xiaomi is pricing MiMo-V2.6-Flash aggressively — at $0.14/$0.28 per 1M tokens, it undercuts most Western frontier-lab offerings by a wide margin while still claiming performance suitable for agentic coding and multi-step tool use. The 15B active-parameter design keeps inference costs low despite the 309B total parameter count, a pattern now common among Chinese labs (DeepSeek, Alibaba Qwen, Moonshot AI) competing on cost-per-token rather than raw parameter count alone.

No independent benchmark verification accompanies this release, and Xiaomi has not published a technical report with specific scores for MiMo-V2.6-Flash, unlike claims made for the earlier V2-Flash version. Buyers evaluating the model for production agentic workloads should treat the "top-tier" and "robust generalization" language as vendor claims pending third-party evaluation.

Related Articles

model release

Xiaomi Launches MiMo-V2.6-Pro-UltraSpeed: Same Quality, 10x Faster Output

Xiaomi's MiMo-V2.6-Pro-UltraSpeed is a fast-inference edition of the company's 1T-parameter flagship MiMo-V2.6-Pro, delivering roughly 10x the output speed at matching quality. It retains the 1M-token context window and native multimodal capabilities, priced at $4.35/$8.70 per 1M input/output tokens.

model release

Xiaomi Launches MiMo-V2.6-Pro, a 1T+ Parameter Model with 1M-Token Context

Xiaomi has released MiMo-V2.6-Pro, a flagship foundation model exceeding 1 trillion parameters with a 1M-token context window and native multimodal support. The model is priced at $0.435 per 1M input tokens and $0.87 per 1M output tokens, targeting agentic and long-horizon tasks.

model release

Yandex Releases AliceAI-Foundation-80B-A3B-Base, an 80B-Parameter MoE Model with 262K Context

Yandex has released AliceAI-Foundation-80B-A3B-Base, an 80-billion-parameter hybrid MoE base model with 3 billion active parameters per token and a 262,144-token context window. The model was trained fully from scratch and, according to Yandex, outperforms larger open-source models on Russian-language factual and educational benchmarks.

model release

xAI Ships Grok 4.7, Cuts Price to $1.60/$4.80 per 1M Tokens With 500K Context

xAI has released Grok 4.7, the successor to Grok 4.6, listed on OpenRouter with a 500K token context window and pricing of $1.60 per 1M input tokens and $4.80 per 1M output tokens. The company claims improvements in long-running software engineering tasks, self-verification, and professional document drafting.

Comments

Loading...