Drooid Logo
Back to story perspectives

Full Breakdown

OpenAI Unveils Jalapeño, a Custom Inference Chip Claiming Speed and Efficiency Gains

8/26/2026, 2:15:11 AM

Core Event

OpenAI introduced its first custom AI inference processor, codenamed Jalapeño, at the Hot Chips conference on August 25, 2026. The ASIC, developed in partnership with Broadcom, is designed to run trained large-language models (LLMs) for tasks such as ChatGPT responses, focusing on lower latency and higher throughput per unit of power.

Background & Context

OpenAI has relied on Nvidia GPUs for most of its compute needs since its inception. Seeking to reduce dependence on general-purpose GPUs and cut inference costs, the company announced the project last year and moved the design to tape-out within nine months. Jalapeño is the first of an intended multigenerational platform that will co-develop chips, memory, and models as an integrated stack.

Data & Statistics

  • Work per watt: OpenAI reports Jalapeño delivers 1.5 – 1.9 × more AI work per watt than Nvidia’s GB200/GB300 superchips across three benchmarked models.
  • Latency: End-to-end latency is 1.7 – 3.6 × lower, with interactive agent workloads showing up to 4.1 × improvement.
  • Power envelope: Benchmark tests were run at 700 watts, measuring throughput per kilowatt and token-per-user rates.
  • Models evaluated: GPT-OSS 120B, DeepSeek R1, Kimi K2.5 1T, and a smaller open-source model.

Official Statements & Responses

The company clarified that Jalapeño is intended solely for inference; training workloads will continue to rely on existing GPU partners, including Nvidia.

Verbatim Quotes

  • “The bottom line is that the results show a very, very significant performance advance over state of the art,” — Richard Ho, hardware vice president

Timeline

  • August 25, 2026: Early proof and benchmark results presented at Hot Chips conference.
  • 2027: Planned scale-up of production and introduction of second-generation designs.

Why It Matters

Inference accounts for the majority of operational cost in AI services, as each user query triggers a full model run. By increasing work per watt and cutting response latency, Jalapeño could lower data-center electricity expenses and improve user experience, potentially enabling lower pricing for AI-powered products. The chip also signals a broader industry shift toward custom silicon for AI workloads, joining efforts by Google (TPU), Amazon, and Anthropic.

What's Next

OpenAI expects the first batch of Jalapeño chips to ship in limited quantities by the end of 2026, with a more substantial rollout in 2027. Ongoing development of second- and third-generation versions is already underway, and the company will continue to benchmark against newer Nvidia architectures as they become available.