OpenAI's Jalapeño Chip Beats Nvidia on Power Efficiency

The Core · TL;DR
- OpenAI revealed benchmark data for its Jalapeño inference chip, built with Broadcom, at the Hot Chips conference.
- The ASIC delivers 1.5-1.9x more throughput per kilowatt and up to 3.6x lower latency than Nvidia's GB200/GB300 systems.
- Small-volume deployment is planned by end of 2026, ramping into 2027, with second and third generations already planned.
- OpenAI says Jalapeño complements rather than replaces its Nvidia-based compute strategy.
OpenAI put numbers behind its custom silicon ambitions on Tuesday, sharing benchmark results for Jalapeño, an inference-focused chip built with Broadcom, at the Hot Chips conference. Richard Ho, OpenAI's head of hardware, led the briefing that laid out how the chip stacks up against Nvidia's current top-end systems.
The results, measured on the InferenceX platform built by analysis firm SemiAnalysis, show Jalapeño delivering 1.5 to 1.9 times more throughput per kilowatt than Nvidia's GB200 and GB300 Blackwell systems. On latency, the gap widens further: across GPT-OSS 120B, DeepSeek R1, and the 1-trillion-parameter Kimi K2.5 model, Jalapeño cut end-to-end response times by 1.7 to 3.6 times.
Jalapeño is an application-specific integrated circuit, meaning it's engineered for one job rather than general-purpose flexibility. OpenAI designed it specifically to speed up the prefill and communication stages of inference, the steps where a model processes an incoming prompt before it starts generating a response, by cutting down on data movement between components.
That focus on inference, rather than training, reflects where OpenAI's costs are increasingly concentrated as its models serve growing numbers of paying users and API calls. Efficiency gains at this stage translate directly into lower serving costs at scale.
Timeline and what comes next
Accounts of when Jalapeño first became public differ slightly. TechCrunch places the initial disclosure in October 2025, while The Verge points to an earlier introduction in June of that year, likely reflecting the difference between an early internal or partner reveal and a wider public announcement. Either way, Tuesday's benchmarks are the first detailed performance data OpenAI has released.
OpenAI plans to deploy Jalapeño in limited volumes by the end of 2026, with broader production ramping through 2027. The company has already committed to building second and third generations of the chip, suggesting this is the start of a multi-year hardware roadmap rather than a one-off project.
OpenAI has been explicit that Jalapeño won't replace its existing chip lineup, and that Nvidia remains part of its compute strategy going forward.
That caveat matters. Custom silicon lets OpenAI tune hardware precisely to its own model architectures and potentially negotiate better economics, but building and scaling chip production carries its own risks and lead times. For now, Jalapeño looks like a complement to Nvidia GPUs rather than a replacement, giving OpenAI more leverage in a market where GPU supply and pricing remain tightly controlled by one dominant vendor.
Original reporting and research used to synthesize this article.
WAKIB Editorial Team
This review was prepared and summarized by the WAKIB AI intelligence engine and vetted by our editorial board for accuracy and reliability.
Subscribe to Newsletter
Get a weekly summary of the most promising AI research and tools delivered to your inbox.
Telegram Channel
Join our active community on Telegram for real-time tracking of AI models and trends.
