benchmarkNVIDIA

Nvidia Claims Groq 3 LPX Hits 3,400 Tokens/Sec, 4x Cerebras — But Needs 64 Chips to Do It

TL;DR

Nvidia's new Groq 3 LPX inference accelerator hit 3,400 tokens per second on Gemma 4 31B, a figure the company says is four times faster than Cerebras. Experts note the benchmark uses at least 64 LPX chips versus Cerebras' one or two accelerators, making the comparison far less clean than it appears.

3 min read
0

Nvidia has moved its specialized inference chip, the Groq 3 LPX, into full production, and an independent benchmark shows record-setting token generation speeds. But the headline comparison against rival Cerebras obscures a major difference in hardware scale, according to analysis from The Register.

At the Hot Chips 2026 conference, Nvidia announced that the Groq 3 LPX has entered full production. The chip, which Nvidia calls an "interactive AI inference accelerator," extends the company's Vera Rubin platform and is designed for fast token generation in agentic AI systems. Nvidia says the chip will go live later this year.

The accelerator is the product of Nvidia's roughly $20 billion acquisition of Groq's license in late December, a deal that also brought Groq founder Jonathan Ross and president Sunny Madra to Nvidia. Groq has historically built processors optimized for inference rather than training — a focus that matters increasingly for agentic applications, which can burn through thousands of tokens across hundreds or thousands of inference steps. Faster token generation means more reasoning steps, tool calls, and verification cycles fit into a user's acceptable wait time. Nvidia claims this can cut coding tasks down to "minutes instead of hours."

The benchmark numbers

A benchmark from Artificial Analysis measured the Groq 3 LPX running the open model Gemma 4 31B with a 100,000-token context window. Across 50 back-to-back requests, the LPX rack hit 3,400 tokens per second, with performance holding steady between 10,000 and 100,000 tokens of input length — the highest figure ever recorded for this model, according to Nvidia. The company says that makes the LPX four times faster than the next-best result: Cerebras, at 882 tokens per second.

Why the comparison is more complicated

The Groq architecture relies on SRAM-heavy dataflow chips, but each LPU carries just 500 MB of memory — 576 times less than a single Rubin GPU's 288 GB, according to The Register. As a result, models must be split across multiple accelerators connected over Ethernet, with a single rack holding up to 256 LPUs. In Nvidia's mixed setup, GPUs handle the compute-heavy prefill phase while LPUs handle the bandwidth-heavy decode phase.

The Register notes that Gemma 4 31B, a dense model that fits entirely within one rack, represents a best-case scenario for this architecture. How it scales to larger mixture-of-experts models is unresolved: DeepSeek V3, for example, would require 1,342 accelerators — just over five racks — to run.

Critically, Nvidia's comparison against Cerebras doesn't account for chip count. Cerebras needs only one or two accelerators to run the same model, while Nvidia's setup requires at least 64 LPX chips to hit its benchmark figure. The comparison also excludes Cerebras' newest CS-4 generation entirely.

Nebius plans to be the first cloud provider to offer the Groq 3 LPX through its Token Factory service, with Groq itself listed among early users.

What this means

Raw tokens-per-second numbers make for a clean marketing claim, but they hide the real cost of the comparison: Nvidia's result depends on scaling out dozens of chips, while Cerebras achieves a competitive figure with a fraction of the hardware. For buyers, the more relevant metric isn't peak throughput but tokens generated per dollar and per chip — figures that aren't yet public. Until Nvidia publishes results against Cerebras' CS-4 or provides efficiency-normalized numbers, the "4x faster" claim should be read as true only under a specific, favorable configuration.

Related Articles

benchmark

NVIDIA's fine-tuned Nemotron 3 Ultra scores 535.4/600 at IOI 2026 and 30/42 at IMO 2026, both above gold

NVIDIA reports that specialized versions of Nemotron 3 Ultra reached gold-medal level at both IOI 2026 (535.4/600) and IMO 2026 (30/42). The IOI run was unofficial, and the IMO proofs were graded by official IMO graders, according to NVIDIA. Checkpoints, datasets, and inference pipelines are published on Hugging Face.

benchmark

GitHub launches ReviewBench, an open benchmark for AI code review agents built on real pull requests

GitHub has launched ReviewBench, an open benchmark for evaluating AI code review agents. According to GitHub, it uses representative pull requests, multi-source ground truth, calibrated evaluation, and production-aligned metrics. Dataset size, model scores, and licensing details were not included in the announcement summary.

benchmark

Aleph Alpha benchmark: Chinese AI models balanced on just 17-41% of 967 sensitive-topic prompts

An Aleph Alpha study of 967 politically sensitive prompts found that only 17 to 41 percent of responses from Alibaba's Qwen, DeepSeek and Moonshot's Kimi were balanced, according to the company's own AI scorer. DeepSeek V4 Pro refused roughly two-thirds of questions. Aleph Alpha sells "sovereign AI" to governments, which gives it a commercial interest in the result.

benchmark

Mercor study: AI models beat 12 licensed CPAs on simplified accounting tasks, but top score on full APEX benchmark is 61

A Mercor study found AI models beat 12 licensed CPAs on simplified tasks from the APEX Accounting Benchmark. On the full 160-task benchmark, the top model, Claude Opus 5.5, meets 61.8% of grading criteria, and no model fully solved almost 60% of tasks.

Comments

Loading...