OpenAI's Jalapeño Inference Chip Beats Nvidia Blackwell on Throughput and Power Efficiency, Company Claims
OpenAI shared the first benchmark results for Jalapeño, its custom inference chip built with Broadcom, at the Hot Chips conference. The company claims the chip beats current Nvidia Blackwell systems on both speed and power efficiency, with limited deployment starting at the end of 2026.
OpenAI disclosed the first benchmark results for Jalapeño, its custom-built inference chip, at the Hot Chips conference on Tuesday. According to OpenAI, the chip outperforms current state-of-the-art inference processors on two key metrics: tokens delivered per user and throughput per kilowatt of power consumed.
The results come from SemiAnalysis's InferenceX benchmark, an independent testing suite for inference hardware. OpenAI compared Jalapeño's performance against an Nvidia Blackwell system — the current commercial standard for AI inference — though the company did not publish specific numeric scores for either system in its presentation.
"The bottom line is that the results show a very, very significant performance advance over state of the art," said Richard Ho, OpenAI's head of hardware, on a press call. "Jalapeño can serve more AI work per unit of power, while also returning responses more quickly. It's very efficient to serve a lot of customers, but it can also be very low latency."
Ho cautioned that the Blackwell comparison reflects current hardware, and that Nvidia's next-generation chips may narrow or erase the gap before Jalapeño reaches meaningful production volume. OpenAI expects to deploy Jalapeño "in very small volumes" by the end of 2026, with broader rollout planned for 2027.
Built with Broadcom, tuned by OpenAI's own models
OpenAI first announced Jalapeño in October, developing the chip in close collaboration with Broadcom. According to OpenAI, its own AI models were used in the chip's design process — part of a broader strategy to treat models, chips, memory, and products as a single co-developed platform rather than separate procurement decisions.
That full-stack approach, OpenAI says, let engineers target specific bottlenecks in the inference pipeline that generic hardware struggles with. The company singled out the prefill and communication phases of inference — the steps where a model processes an incoming prompt and moves data between compute units — as frequent sources of latency.
"We designed Jalapeño to minimize data movement and communication delays," OpenAI said in a blog post accompanying the benchmark release. "This means that model state, including the KV cache used while generating a response, can be explicitly placed and kept local while the system activates the right combination of compute, memory, and networking for each inference phase."
Pricing, exact chip specifications, and node process details have not been disclosed.
What this means
Jalapeño is OpenAI's clearest move yet toward reducing its dependence on Nvidia for the compute that powers ChatGPT and its API. Inference — not training — is where OpenAI now spends the bulk of its compute budget, since every user query requires running the model, while training happens only periodically. Shaving latency and power costs at that scale compounds quickly across hundreds of millions of daily requests.
The claims should be read with appropriate caution: OpenAI benchmarked against Blackwell, a chip Nvidia has already begun to supersede, and the company published no absolute performance numbers, only comparative claims verified through a third-party benchmark it selected. Real validation will come once Jalapeño ships in volume in 2027 and independent operators can test it under production workloads. Until then, this is a proof of concept from a company that has strong incentive to prove its custom silicon bet is paying off — and an early signal that the largest AI labs are no longer willing to treat inference hardware as someone else's problem.
Related Articles
Mathematicians' group calls for OpenAI boycott after release of 700+ AI-generated proof files
The Association for Human Mathematics (AHM), chaired by Fields Medalist Terence Tao, is urging mathematicians to stop working with OpenAI after the company released more than 700 AI-generated manuscripts at once. The group says the release violates scientific norms. Critics say many of the papers are too dense to verify without AI assistance.
Only 10 of OpenAI's 719 math manuscripts include chain of thought, falling short of expert guidelines
OpenAI released 719 manuscripts claiming solutions to open math problems, but only 10 include the model's chain of thought. A Cambridge and King's College London paper also documents at least two discrepancies between the natural-language and Lean versions of OpenAI's Navier-Stokes-derived result.
OpenAI launches GPT-6 in ChatGPT with 'Intelligent UI' and interactive answers; Sol for paid users, Luna for free
OpenAI is rolling out GPT-6 to all ChatGPT tiers, with paying users on GPT-6 Sol and free users on GPT-6 Luna. The release adds 'Intelligent UI,' which renders answers as interactive charts, buttons, forms and mini apps, and lets the model respond while still thinking. OpenAI claims this cuts wait times by 44 percent.
Common Sense Media rates ChatGPT for Teens an 'unacceptable risk,' citing failures on 3 of 5 Red Lines
Common Sense Media has labeled OpenAI's ChatGPT for Teens an "unacceptable risk," saying it failed three of five severe-harm Red Lines and kept using engagement cues during crisis conversations. OpenAI disputes the testing methodology, saying testing may have ended before parental controls were fully active.
Comments
Loading...