OpenAI's Jalapeño Inference Chip Beats Nvidia Blackwell on Throughput and Power Efficiency, Company Claims
OpenAI shared the first benchmark results for Jalapeño, its custom inference chip built with Broadcom, at the Hot Chips conference. The company claims the chip beats current Nvidia Blackwell systems on both speed and power efficiency, with limited deployment starting at the end of 2026.
OpenAI disclosed the first benchmark results for Jalapeño, its custom-built inference chip, at the Hot Chips conference on Tuesday. According to OpenAI, the chip outperforms current state-of-the-art inference processors on two key metrics: tokens delivered per user and throughput per kilowatt of power consumed.
The results come from SemiAnalysis's InferenceX benchmark, an independent testing suite for inference hardware. OpenAI compared Jalapeño's performance against an Nvidia Blackwell system — the current commercial standard for AI inference — though the company did not publish specific numeric scores for either system in its presentation.
"The bottom line is that the results show a very, very significant performance advance over state of the art," said Richard Ho, OpenAI's head of hardware, on a press call. "Jalapeño can serve more AI work per unit of power, while also returning responses more quickly. It's very efficient to serve a lot of customers, but it can also be very low latency."
Ho cautioned that the Blackwell comparison reflects current hardware, and that Nvidia's next-generation chips may narrow or erase the gap before Jalapeño reaches meaningful production volume. OpenAI expects to deploy Jalapeño "in very small volumes" by the end of 2026, with broader rollout planned for 2027.
Built with Broadcom, tuned by OpenAI's own models
OpenAI first announced Jalapeño in October, developing the chip in close collaboration with Broadcom. According to OpenAI, its own AI models were used in the chip's design process — part of a broader strategy to treat models, chips, memory, and products as a single co-developed platform rather than separate procurement decisions.
That full-stack approach, OpenAI says, let engineers target specific bottlenecks in the inference pipeline that generic hardware struggles with. The company singled out the prefill and communication phases of inference — the steps where a model processes an incoming prompt and moves data between compute units — as frequent sources of latency.
"We designed Jalapeño to minimize data movement and communication delays," OpenAI said in a blog post accompanying the benchmark release. "This means that model state, including the KV cache used while generating a response, can be explicitly placed and kept local while the system activates the right combination of compute, memory, and networking for each inference phase."
Pricing, exact chip specifications, and node process details have not been disclosed.
What this means
Jalapeño is OpenAI's clearest move yet toward reducing its dependence on Nvidia for the compute that powers ChatGPT and its API. Inference — not training — is where OpenAI now spends the bulk of its compute budget, since every user query requires running the model, while training happens only periodically. Shaving latency and power costs at that scale compounds quickly across hundreds of millions of daily requests.
The claims should be read with appropriate caution: OpenAI benchmarked against Blackwell, a chip Nvidia has already begun to supersede, and the company published no absolute performance numbers, only comparative claims verified through a third-party benchmark it selected. Real validation will come once Jalapeño ships in volume in 2027 and independent operators can test it under production workloads. Until then, this is a proof of concept from a company that has strong incentive to prove its custom silicon bet is paying off — and an early signal that the largest AI labs are no longer willing to treat inference hardware as someone else's problem.
Related Articles
OpenAI Reinstates 5-Hour Usage Limit for ChatGPT Plus Codex and Work Tiers
OpenAI will reinstate a five-hour usage limit for Codex and ChatGPT Work on Plus subscriptions starting August 25, 2026, after weeks of running only a weekly cap. Pro $100 and Pro $200 plans remain exempt from the change for the foreseeable future.
OpenAI Launches ChatGPT Plugin That Reads and Analyzes Mac iMessages
OpenAI released a new ChatGPT plugin for Mac that connects to Apple's Messages app, letting the AI send texts, search and summarize conversations, and analyze communication patterns with specific contacts. The feature requires macOS permission grants to read message content.
OpenAI Launches ChatGPT Plugin That Reads and Replies to iMessages on Mac
OpenAI has released a plugin allowing ChatGPT to read and respond to Apple iMessages on Mac, currently limited to ChatGPT Work and Codex users. The feature requires users to grant Full Disk Access and contact permissions, and arrives amid an active lawsuit between Apple and OpenAI.
OpenAI Adds Transparent Background Generation to GPT-Image-2 API
OpenAI is previewing a transparent background feature for GPT-Image-2 through its API, letting developers generate PNGs with no background baked in at generation time. The company claims this produces cleaner results than traditional background removal, particularly on difficult edges like glass or thin fibers.
Comments
Loading...