First benchmark results for OpenAI’s custom Jalapeño inference chip report up to 1.9x more AI work per watt and as much as 3.6x lower latency than leading Nvidia accelerators.
Picture asking ChatGPT a complex question and getting the answer in just over a second. Then picture the same question taking nearly two seconds on the hardware that powered it yesterday. That gap — roughly 0.7 to 0.8 seconds — is what OpenAI is pointing to as proof that its first custom-built chip, named Jalapeño, step forward in how AI responses are delivered at scale.
In August 2026, OpenAI published its first measured benchmark results for Jalapeño, the custom inference chip it announced in June 2026 in partnership with Broadcom. The numbers it posted are striking, though it’s worth keeping in mind they come from OpenAI’s own test environments rather than a neutral laboratory.
What Jalapeño Actually Is
OpenAI calls Jalapeño an “Intelligence Processor” — its own term for what is, technically, an AI inference accelerator. The distinction matters. Inference chips are built for *running* AI models, not training them. When you type a prompt into ChatGPT and receive a reply, that’s inference. Training — the process of teaching the model in the first place — is a different, enormously power-hungry workload that still largely depends on Nvidia hardware.
Jalapeño is designed specifically for the inference side, and its architecture reflects that focus. Rather than splitting the work across multiple chips or heterogeneous configurations, it handles both phases of large language model inference in a single design: the “prefill” phase, which processes your prompt, and the “decode” phase, which generates each word of the reply. Keeping both on one chip, OpenAI says, reduces the synchronisation overhead that slows down multi-chip systems.
Broadcom manufactures the silicon. Each accelerator reportedly runs at around 700 watts — considerably less than the 1,200 to 1,400 watts associated with Nvidia’s GB200 and GB300-class systems. That lower power draw is central to OpenAI’s efficiency claims.
The Numbers OpenAI Is Putting Forward
The benchmarks use InferenceX, a publicly documented test suite maintained by SemiAnalysis, run across three large models: GPT-OSS-120B, DeepSeek R1, and Kimi K2.x.
On the GPT-OSS-120B model, Jalapeño-based systems reportedly reached around 85,000 mixed tokens per second per kilowatt. The Nvidia GB200 baseline managed about 45,000. That’s roughly 1.9 times more AI work for every unit of electricity consumed.
Latency improvements are even sharper in some tests. End-to-end response time on GPT-OSS-120B dropped from about 1.8 seconds on the Nvidia baseline to around 1.0 to 1.1 seconds on Jalapeño. The time between individual tokens — the rhythm at which words appear on screen — fell from about 1.9 milliseconds to around 0.7 milliseconds, enabling per-user generation speeds in the region of 1,400 to 1,500 tokens per second. Across all tested models, OpenAI reports efficiency gains of 1.5x to 1.9x per kilowatt and latency reductions of 1.7x to 3.6x. For DeepSeek R1 and Kimi K2.x, some tests indicated more than three times faster end-to-end response at comparable throughput.
At the system level, one described Jalapeño configuration packs 128 accelerators into a single rack, delivering roughly 1.7 exaFLOPS of 4-bit compute, around 27.5 terabytes of HBM4 memory, and close to 2 petabytes per second of memory bandwidth. Individual accelerators are reported — partly by inference from published specifications rather than direct OpenAI disclosure — to offer around 13.4 petaFLOPS at 4-bit precision and roughly 15 terabytes per second of memory bandwidth.
How Independent Analysts View the Claims
Not everyone is simply taking the numbers at face value.
Dylan Patel, founder of SemiAnalysis, whose InferenceX benchmark suite was used in the tests, has discussed the methodology publicly, and independent analysts broadly treat the results as credible — but vendor-driven. The tests use a public suite and publish some configuration details. But the headline numbers still originate from OpenAI-controlled experiments. Calls for independent replication and fuller disclosure of software stacks and exact system configurations are already circulating in the AI hardware community.
There’s also a broader concern. Jalapeño may reduce OpenAI’s dependence on Nvidia at the hardware level, but it deepens vertical integration in a different direction — one where OpenAI controls the model, the chip, the runtime software, and the data centre interconnect. The UK’s Competition and Markets Authority and Ofcom have both been examining concentration of power among large AI and cloud providers, and proprietary silicon of this kind is likely to feature in future discussions about interoperability and market structure, even if Jalapeño hasn’t been named in any regulatory document yet.
And there’s the Jevons paradox to consider. A more efficient chip doesn’t necessarily mean lower overall energy consumption — it can simply enable more workloads to run, pushing total electricity use upward. Environmental groups monitoring AI’s energy footprint have raised this concern repeatedly as efficiency improvements are announced across the sector.
OpenAI’s Broader Strategy
Jalapeño isn’t a one-off experiment. OpenAI describes it as the first chip in a planned multi-generation compute platform, part of a “full stack” strategy in which models, software, interconnects, and silicon are all designed together rather than bought separately. Google has its TPU line. Amazon has Trainium and Inferentia. Meta is building its own accelerators. OpenAI, one of the last of the major AI developers to enter this space, is now firmly in the race.
For now, Jalapeño is being deployed in OpenAI’s own data centres to power services including ChatGPT and its developer APIs. There is no confirmed public retail product. OpenAI has not published any pricing changes tied explicitly to the chip.
What This Means for Kent Residents
For people and organisations in Kent using ChatGPT or tools built on OpenAI’s API — from NHS Kent and Medway ICB exploring administrative automation to local businesses using AI-assisted software — the most visible effect of Jalapeño, if OpenAI routes more inference workloads onto the new hardware, would be faster response times for complex queries. OpenAI has not announced any pricing changes linked to the chip, so cost benefits remain speculative for now. More broadly, efficiency gains of the kind Jalapeño claims — nearly half the power draw per unit of useful AI work compared with current Nvidia systems — are relevant to UK energy and planning debates around data centre expansion, as local authorities across the South East weigh the infrastructure demands of large-scale AI computing against grid capacity and sustainability targets.
Source: @OpenAI
OpenAI's Jalapeño Chip Claims Up to 1.9x Better Efficiency Than Nvidia's Best AI Hardware Quiz
5 questions