Inception
https://www.inceptionlabs.ai →News
No articles yet.
Models
Mercury 2
Inception
The first reasoning-capable diffusion LLM. Hits 1,009 tokens/sec on NVIDIA Blackwell with 1.7s end-to-end latency — 5x faster than speed-optimized autoregressive models. Tunable reasoning, native tool use, schema-aligned JSON output. OpenAI-compatible API.
Context128K
Input/1M$0.25
Feb 24, 2026
Mercury
Inception
The first commercial-scale diffusion large language model (dLLM). Generates text by refining many tokens in parallel instead of one at a time, reaching 1,000+ tokens/sec. Now legacy — superseded by Mercury 2.
Context32K
Jun 26, 2025