Drooid Logo
Back to story perspectives

Full Breakdown

OpenAI Unveils Jalapeño, Its First Custom AI Inference Chip

6/24/2026, 10:52:56 PM

Core Event

On June 24 2026, OpenAI announced Jalapeño, a custom ASIC built with Broadcom for inference workloads. The chip will run large-language models such as ChatGPT, Codex and GPT-5.3-Codex-Spark in OpenAI’s data centers.

Background & Timeline

OpenAI’s reliance on Nvidia GPUs grew costly during a global GPU shortage. In October 2025 the company disclosed a partnership with Broadcom to design its own accelerator, mirroring similar moves by Google, Amazon, Microsoft and Meta. Using its own models to aid design, OpenAI completed the chip’s tape-out in nine months—a cycle typically lasting 18-24 months. Engineering samples were delivered on June 24 2026. Small-scale prototype deployments are slated for late 2026, with full production and gigawatt-scale data-center integration planned for 2027-2028.

Impact & Data

Early lab tests show performance-per-watt “substantially better” than current GPUs, while Broadcom CEO Hock Tan cites roughly 50 % lower inference cost per token. The chip’s rapid development and the nine-month design window are highlighted as record-fast for high-performance semiconductors. Broadcom’s share price rose 10 % on the announcement, up about 7 × since late 2022. OpenAI targets 10-13 GW of Jalapeño-powered compute by 2029.

Responses & Criticism

OpenAI President Greg Brockman described the chip as part of a “full-stack infrastructure strategy” that leverages a deep workload understanding to make compute more affordable. Hardware lead Richard Ho said the design is optimized for memory, kernel and networking patterns of frontier models. Broadcom’s Hock Tan emphasized “insatiable” compute demand and claimed performance parity with Nvidia’s Blackwell GPUs and Google’s TPUs. Analysts note the 50 % cost-saving claim lacks independent verification, and OpenAI reported an operating loss of nearly $20.92 billion in 2025, driven largely by compute expenses. The company also confirmed continued reliance on Nvidia GPUs for model training.

Conflicting Reports & Gaps

Broadcom quantifies a 50 % reduction in inference cost, whereas OpenAI’s public statement only promises “substantially better” performance-per-watt without a specific figure. Neither party has released detailed benchmark data or disclosed the comparison baseline, leaving the true magnitude of the advantage unconfirmed.

Verbatim Quotes

  • “We have a deep understanding of the workload,” — Greg Brockman, President, OpenAI
  • “The degree to which our models have been able to accelerate it was very surprising to us,” — Greg Brockman, President, OpenAI (CNBC interview)
  • “Based on early testing, Jalapeño will efficiently execute our most important workloads close to the hardware’s theoretical limits,” — Richard Ho, Head of Hardware, OpenAI
  • “Broadcom CEO Hock Tan said the accelerator is showing roughly 50% cost savings compared with typical AI graphics processing units.” — Hock Tan, CEO, Broadcom (TradingView)

What’s Next

OpenAI will begin limited prototype deployments of Jalapeño by the end of 2026, scale production through 2027-2028, and launch a second-generation chip in 2028. The partnership aims to integrate the processors into gigawatt-scale data centers with Microsoft and other cloud partners.