OpenAI and Broadcom Unveil Jalapeño, OpenAI’s First Custom AI Chip for LLM Inference

OpenAI and Broadcom Unveil Jalapeño, OpenAI's First Custom AI Chip for LLM Inference

OpenAI says Jalapeño is its first AI chip, built with Broadcom for large language model inference across ChatGPT, Codex, the API and future agentic products.

OpenAI has announced Jalapeño, its first custom silicon accelerator, developed in partnership with semiconductor giant Broadcom. The chip is designed specifically for large language model inference — the process of running AI models to generate responses — rather than for training new models from scratch.

The announcement positions Jalapeño as the first in a planned multi-generation compute platform, with OpenAI describing it as its first Intelligence Processor. According to OpenAI, the chip went from design to production in nine months, a timeline the company presents as a marker of how quickly it can now move in hardware development.

What Jalapeño Is — and What It Is Not

The chip is not a general-purpose AI accelerator adapted from an earlier design. OpenAI says it was built from a blank slate, shaped entirely around the workloads it actually runs: ChatGPT, Codex, its developer API, and the agentic products it is building towards. That distinction matters. Most AI chips in use today were designed with training workloads in mind, then pressed into inference service. Jalapeño, according to OpenAI, starts from the opposite assumption.

Broadcom’s role in the partnership covers silicon implementation alongside networking and connectivity technologies within the joint platform. OpenAI’s announcement frames the collaboration as a full-stack effort, not simply a manufacturing contract.

Early internal testing, according to OpenAI, shows performance per watt substantially better than what the company describes as the current state of the art. But no final technical performance data has been published. Independent benchmarking is not yet available, so those claims cannot be verified externally.

The Scale of What OpenAI Is Planning

The ambition here is large. OpenAI says the first-generation accelerator is intended for deployment at gigawatt scale with data-centre partners across multiple generations of the platform. To put that in context, a single gigawatt of power capacity is roughly equivalent to the output of a large nuclear power station. Running AI inference at that scale would represent a hefty slice of global data-centre energy consumption.

That energy dimension is not incidental. Technology media covering the announcement have noted that inference efficiency — getting more output per unit of electricity — has become one of the central cost problems in the AI industry. Serving hundreds of millions of users through products like ChatGPT requires enormous and continuous compute. Chips that do that job more efficiently reduce the cost per query.

Sam Altman, OpenAI’s chief executive, said: “Chips are foundational to AI.”

It’s a short statement. But it reflects a strategic shift that has been building for some time.

Reducing Dependence on NVIDIA

OpenAI has not officially stated that Jalapeño is intended to reduce its dependence on NVIDIA. But technology analysts and media commentators have framed it that way, and the logic is straightforward. NVIDIA currently dominates the market for AI accelerators, and its H100 and B200 GPUs are the hardware on which most large language models — including OpenAI’s own — are currently trained and served. Building proprietary inference silicon gives OpenAI more control over its own supply chain, its cost base, and its ability to scale without competing for chips in a constrained market.

OpenAI’s announcement describes Jalapeño as part of a broader expansion from products and models into chips — what it calls building out a full-stack platform. That framing suggests the company sees hardware not as a commodity input but as a layer of its business it wants to own.

One figure circulating in secondary coverage claims the chip could deliver up to 50 per cent cost savings on inference. That figure is unverified and does not appear in OpenAI’s official announcement materials; it should be treated with caution until the company publishes confirmed data.

What Comes Next

Several questions remain open. OpenAI has not published a deployment timeline for Jalapeño beyond the gigawatt-scale ambition. It has not confirmed which data-centre partners will host the hardware, nor when external developers using the API might see any change in performance or pricing. The company has also not said when or whether it will share detailed technical specifications.

The chip’s existence is confirmed. Whether it performs as described at production scale is a question that will only be answered once it is running at volume.

What This Means for Kent Residents

For Kent residents who use ChatGPT, OpenAI’s API, or AI-powered tools built on top of OpenAI’s services, the long-term promise of Jalapeño is faster and potentially cheaper AI inference — though no pricing changes have been announced and no timeline has been confirmed. Businesses across the county that rely on AI services for productivity, customer support, or software development may eventually benefit if the chip delivers the efficiency gains OpenAI is claiming, but that outcome depends entirely on performance data that has yet to be published.

Source: @OpenAI

OpenAI and Broadcom Unveil Jalapeño, OpenAI's First Custom AI Chip for LLM Inference Quiz

5 questions