Inception Launches Mercury 2.5 Preview, a Diffusion LLM Claiming 1,107 Tokens/Sec
Inception released Mercury 2.5 Preview, a diffusion-based language model that generates tokens in parallel rather than sequentially, claiming throughput of 1,107 tokens per second on standard GPUs. The model is available on OpenRouter with a 260K context window and an 80% launch discount through September 8, 2026.
Mercury 2.5 Preview — Quick Specs
Inception Releases Mercury 2.5 Preview
Inception has released Mercury 2.5 Preview, the latest version of its diffusion large language model (dLLM), available now through OpenRouter. Unlike standard autoregressive models that generate tokens one at a time, Mercury 2.5 produces and refines multiple tokens in parallel, a diffusion-based approach the company says makes it the fastest reasoning LLM currently available.
Key specs
- Context window: 260,000 tokens
- Throughput: 1,107 tokens/sec on standard GPUs, according to Inception
- P50 latency: 1.38 seconds, with observed throughput of 223 tokens/sec on OpenRouter's tracked provider endpoint
- Pricing: $0.20 per 1M input tokens / $0.75 per 1M output tokens (list price); currently $0.04 / $0.15 per 1M tokens under an 80% launch discount running through September 8, 2026 at 07:00 UTC
- Cache read pricing: $0.02 per 1M tokens (discounted from $0.75)
- Release date: August 31, 2026
Claimed performance
Inception claims Mercury 2.5 delivers a "10+ point jump in intelligence" over its predecessor, Mercury 2, though the company did not specify which benchmark suite produced that figure. Inception positions the model's quality as comparable to cost-optimized frontier offerings, specifically naming GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5 as reference points. None of these comparisons have been independently verified.
Features
Mercury 2.5 supports tunable reasoning levels, letting developers dial compute up or down depending on task difficulty. It also supports parallel tool calls and schema-aligned JSON output, features aimed at agentic and structured-output workloads. Inception markets the model for latency-sensitive production use cases: search agents, voice pipelines, and coding subagents, where the cost of sequential token generation compounds across long interaction chains.
Availability data
OpenRouter's tracked deployment shows 100% uptime over the trailing three days but 97.13% availability over the same window, with 24-hour availability at 97.19%. OpenRouter notes that failed requests can be routed to alternate healthy providers if a request's routing filters allow it.
What this means
Mercury 2.5 is notable less for a specific benchmark score and more for its architecture: diffusion-based generation is still a minority approach among production LLMs, and a claimed 1,107 tokens/sec would be a meaningful jump over typical autoregressive throughput on comparable hardware. If Inception's speed claims hold up under independent testing, this positions Mercury 2.5 as a serious option for latency-bound applications like voice agents and multi-step coding subagents where waiting on token-by-token generation is the primary bottleneck.
The comparisons to GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5 suggest Inception is targeting the budget-tier segment of the market rather than competing with top-end frontier models on raw capability. The 80% discount pricing — effectively $0.04/$0.15 per 1M tokens against a $0.20/$0.75 list price — is a temporary promotional window ending September 8, 2026, after which buyers should expect list pricing to apply. As with all vendor-reported throughput and quality claims, third-party benchmarking will be needed to confirm the stated gains over Mercury 2.
Related Articles
Alibaba Releases Qwen3.8 Flash, a Multimodal Reasoning Model with 1M-Token Context
Alibaba has released Qwen3.8 Flash, a multimodal reasoning model with a 1 million token context window, aimed at coding, agentic workflows, and visual/document analysis. It's priced at $0.16 per 1M input tokens and $0.47 per 1M output tokens through Alibaba Cloud International.
IBM Releases Granite 4.2 8B, a Dense Reasoning Model with 131K Context and Three Thinking Modes
IBM has released Granite 4.2 8B, a dense reasoning model built for math, code generation, and agentic workflows. The model supports 131K context, 12 languages, and three switchable reasoning modes, priced at $0.10 per 1M input tokens and $0.15 per 1M output tokens.
Tencent Releases Hy4 Preview: 770B-Parameter MoE Model with 1M Context for Coding Agents
Tencent has released Hy4 preview, a mixture-of-experts model with 770B total parameters and 49B active parameters, targeting coding agents and multi-step tool-use workflows. The model ships with a 1 million token context window and is priced at $0.834 per 1M input tokens and $2.501 per 1M output tokens.
Z.ai Confirmed as Creator of Chart-Topping 'Ox Alpha' Model, Weights Coming Wednesday
Z.ai, maker of the GLM model series, has confirmed it is behind Ox Alpha, the mysterious open-weight model that appeared anonymously on OpenRouter and topped benchmark leaderboards. The company will release the model's weights on Wednesday.
Comments
Loading...