TPSTokens Per Second
NewsModelsIDEsRankingsBenchmarksSecurity

Navigate

NewsModelsIDEsRankingsChangelogBenchmarksSecurityModel QuizWhat I MissedHistorySaved

Companies

AnthropicOpenAIDeepMindMeta AIDeepSeekxAIMistral AIPerplexity AI

Inception

https://www.inceptionlabs.ai →

News

No articles yet.

Models

Mercury 2

Inception

active

The first reasoning-capable diffusion LLM. Hits 1,009 tokens/sec on NVIDIA Blackwell with 1.7s end-to-end latency — 5x faster than speed-optimized autoregressive models. Tunable reasoning, native tool use, schema-aligned JSON output. OpenAI-compatible API.

Context128K
Input/1M$0.25

Feb 24, 2026

Mercury

Inception

deprecated

The first commercial-scale diffusion large language model (dLLM). Generates text by refining many tokens in parallel instead of one at a time, reaching 1,000+ tokens/sec. Now legacy — superseded by Mercury 2.

Context32K

Jun 26, 2025

Top Benchmark Scores

Full leaderboard →

AIME 2025

Mercury 2
91.1%

Arena Elo

Mercury 2
1347 elo

GPQA

Mercury 2
77%

Hallucination Rate

Mercury 2
12.3%

LiveCodeBench

Mercury 2
67%

Speed (tok/s)

Mercury 2
1185 tokens_per_sec
TPS

Tokens Per Second. The fastest LLM news on the internet — tracked automatically every 15 minutes.

Coverage

  • Latest News
  • Model Database
  • AI IDEs
  • Compare IDEs
  • Changelog
  • Benchmarks
  • Compare Models

Companies

  • Anthropic
  • OpenAI
  • Google DeepMind
  • Meta AI
  • DeepSeek
  • xAI
  • Mistral AI
  • Perplexity AI

Guides

  • All Rankings
  • Best by Use Case
  • Best Coding LLM
  • Best Cheap LLM
  • Best Reasoning LLM
  • Best AI IDE
  • Compare Models
  • Compare IDEs
  • About TPS
  • RSS Feed
  • Atom Feed

© 2026 TPS — Tokens Per Second.

The fastest LLM news. All signal, no noise.