Mercury 2
Inception🇺🇸 United States
The first reasoning-capable diffusion LLM. Hits 1,009 tokens/sec on NVIDIA Blackwell with 1.7s end-to-end latency — 5x faster than speed-optimized autoregressive models. Tunable reasoning, native tool use, schema-aligned JSON output. OpenAI-compatible API.
Context window128K tokens
Input / 1M tokens$0.25
Output / 1M tokens$0.75
Version History
mercury-2-launchmajor
Mercury 2 launches. The first reasoning-capable diffusion LLM. Hits 1,009 tokens/sec on NVIDIA Blackwell with 1.7s end-to-end latency — 5x faster than speed-optimized autoregressive models. Tunable reasoning, native tool use, schema-aligned JSON output. OpenAI-compatible API.
Benchmark Scores
Full leaderboard →91.1%
AIME 2025
1347.0 elo
Arena Elo
77.0%
GPQA
12.3%
Hallucination Rate
67.0%
LiveCodeBench
1185.0 tokens_per_sec
Speed (tok/s)