OpenAI and Broadcom Unveil Jalapeño: A Custom LLM Inference Chip Targeting Gigawatt-Scale Deployment

The Core · TL;DR
- OpenAI and Broadcom unveiled Jalapeño on June 24, 2026 — OpenAI's first custom LLM inference accelerator, developed in just nine months.
- Engineering samples are already running real ML workloads including GPT-5.3-Codex-Spark, with early testing showing substantially better performance-per-watt than current state-of-the-art chips.
- Jalapeño is designed to work with all LLMs and is planned for gigawatt-scale deployment across data center partners over multiple generations.
- Initial production deployment is targeted by end of 2026, with OpenAI's hardware program led by Richard Ho.
OpenAI and Broadcom have jointly announced Jalapeño, OpenAI's first proprietary Intelligence Processor — a silicon accelerator purpose-built for large language model inference. The chip was formally unveiled on June 24, 2026, with Broadcom President and CEO Hock Tan and President Charlie Kawwas delivering the first units directly to OpenAI CEO Sam Altman and President Greg Brockman.
Nine Months from Whiteboard to Silicon
What makes the Jalapeño announcement particularly striking is the pace of its development. The chip went from initial design to production-ready hardware in just nine months — a timeline that would be considered aggressive even for established semiconductor players with decades of chip design experience. OpenAI's hardware program, led by Richard Ho, appears to have moved with unusual velocity by partnering with Broadcom rather than attempting to build foundational chip design capabilities in-house from scratch.
Architecture and Performance Claims
Jalapeño is explicitly architected for LLM inference, not training — a design choice that reflects where OpenAI's most pressing computational costs and latency constraints currently sit. Unlike general-purpose GPUs repurposed for AI workloads, the chip is designed from the ground up to run inference across all LLMs, not just OpenAI's own models.
Early benchmark data is promising, though still preliminary: engineering samples are already executing real ML workloads, including GPT-5.3-Codex-Spark, and OpenAI reports that performance-per-watt metrics are substantially ahead of the current state of the art. Specific numerical benchmarks have not yet been published, so independent verification of these claims remains pending.
Deployment Ambitions
OpenAI's stated deployment targets are ambitious. The company plans to roll Jalapeño out at gigawatt-scale across data center partners, with the program spanning multiple chip generations. Initial deployment is targeted before the end of 2026 — a tight window that suggests engineering samples are already close to production-grade quality rather than early prototype silicon.
The gigawatt framing is a deliberate signal: this is not a niche research accelerator or a limited internal deployment. It positions Jalapeño as a potential cornerstone of OpenAI's infrastructure strategy, reducing reliance on third-party GPU suppliers while enabling tighter optimization of the full stack from model architecture down to silicon.
Strategic Context
For Broadcom, the partnership deepens its role as a preferred ASIC design partner for hyperscalers and AI-native companies — a position it has been building alongside relationships with Google (TPU) and others. For OpenAI, owning the inference silicon layer closes a critical gap in its vertical integration story, one that rivals like Google and Amazon have been executing on for years.
Whether Jalapeño delivers on its performance-per-watt promise at scale will be the definitive test — and the data center deployments planned for late 2026 will be the first real proving ground.
Original reporting and research used to synthesize this article.
WAKIB Editorial Team
This review was prepared and summarized by the WAKIB AI intelligence engine and vetted by our editorial board for accuracy and reliability.
Subscribe to Newsletter
Get a weekly summary of the most promising AI research and tools delivered to your inbox.
Telegram Channel
Join our active community on Telegram for real-time tracking of AI models and trends.
