OpenAI has released the first measurements for the Jalapeño Chip, the inference part it is building with Broadcom. On SemiAnalysis’s InferenceX suite, it serves up to 1.9 times more work per kilowatt than Nvidia’s GB200 and GB300 racks, with end to end latency up to 3.6 times lower. Deployment, on the other hand, stays symbolic until 2027.
Key Takeaways
- Jalapeño draws 700 watts, against Nvidia racks rated at 1,200 and 1,400 watts.
- The gain runs from 1.5 to 1.9 times on throughput per kilowatt, and climbs to 4.1 times on interactive workloads.
- OpenAI is promising very small volumes late in 2026, with meaningful capacity arriving in 2027.
Have an AI Sum Up This Article
ChatGPTWhat the Hot Chips numbers actually cover
The chip itself was no secret. OpenAI announced in October 2025 that it was working with Broadcom on silicon cut for its own workloads. What was missing were the figures, and they landed Tuesday at the Hot Chips conference, covering inference only, the phase where a model answers a user rather than learns.
The chosen yardstick is InferenceX, a public suite published by the research shop SemiAnalysis. Its point is to measure the full serving path of a request instead of raw compute, which lands the result much closer to what a user feels on screen.
Three open weight models carried the load. GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T sit at very different sizes, which keeps a verdict from resting on a single memory profile.
The result comes in two series. The Jalapeño Chip delivers 1.5 to 1.9 times more throughput per kilowatt than the GB200 and GB300 racks, and returns answers with end to end latency 1.7 to 3.6 times lower, as OpenAI’s own write-up of the first Jalapeño results lays out.
The power gap is what gives the exercise its weight. The tested part is rated at 700 watts, facing accelerators rated at 1,200 and 1,400 watts. On interactive workloads, the ones where somebody is waiting in front of a screen, the claimed edge widens to 4.1 times.
Richard Ho, who runs hardware at OpenAI, framed it plainly. The chip serves more AI work per unit of power while also returning responses more quickly. Those two properties rarely arrive together, since one is normally bought with the other.
One manufacturing detail is worth flagging. OpenAI says its own models helped design the part, a loop the company has been advertising for a year and that nobody outside can quantify. It also frames the Jalapeño Chip as the first stone of a multigenerational platform meant to line up products, models, silicon and memory rather than as a standalone component.
The architecture explains part of the gain. The Jalapeño Chip cuts data movement, keeps model state and the attention cache local, and tunes the compute, memory and networking mix to whichever inference phase is running. Prefill and chip to chip exchange are the two places where time leaks, and those are the ones that got targeted.
A gain that will not land in any 2026 budget
The calendar cools the announcement straight away. OpenAI talks about very small volumes by the end of the year, and meaningful deployment in 2027. No team can build a capacity plan on this part during the current fiscal year.
The context makes it urgent anyway. Inference demand has changed scale in a matter of months, to the point that agents now burn more tokens than humans on routing platforms. A fleet answering agents runs around the clock, while a fleet answering people gets to breathe at night.
That shift is what gives the watt figure its meaning. Once load becomes permanent, energy stops being one cost line among many and turns into the physical ceiling of the service. A chip that returns twice the work on the same electricity pushes that ceiling back without pouring another slab of concrete.
OpenAI is pouring concrete regardless. The company just closed a $105 billion financing package for its Ohio data center, which shows the in-house chip is not meant to replace real estate but to squeeze more out of it.
The same constraint shows up in internal trade-offs. The company recently froze its largest training run, a reminder that available capacity is split between what learns and what answers. Every watt handed back to inference is a watt that never has to be argued against research.
One caveat still stands. The suite is public, but the runs and the publication both come from the vendor. Until a third party replays the measurement on its own fleet, these ratios remain a documented claim rather than an independent check.
What OpenAI does with the savings is the open question. It can keep them to absorb the growth of its agent traffic, or pass them into pricing to squeeze the labs that rent their capacity. The first option protects margin, the second attacks the market, and Tuesday’s publication settles neither.
More articles on Horizon
- Agent tokens have jumped fourteen times since February
- Wan 3.0 Builds Thirty Seconds of Video From a PDF
- V4-Flash-Vision Test: DeepSeek Wins on Agent Work
Nvidia loses an argument, not a customer
Nvidia’s position does not rest on power efficiency alone. The company sells a software environment thousands of teams already know, and that asset does not get cloned in one silicon generation. Jalapeño will not be sold to third parties either, which keeps it out of the market Nvidia actually fights in.
What does change is the talking point. Until now the standard answer to in-house chips was to point at their performance gap. The kilowatt figure removes that answer on the one segment that dominates the running cost of a consumer service.
The clearest interim winner is Broadcom. Every lab looking to loosen its grip on a single supplier has to go through a partner able to design and industrialize a custom accelerator, and that skill set fits on one hand.
For rival labs the pressure comes from two directions. Those buying capacity at market price now face a competitor whose marginal cost per request is set to fall, and those preparing their own silicon have to publish comparable measurements or read as followers.
The next signal to watch is simple enough. If the first units go live on real traffic late in 2026, and if OpenAI pricing moves shortly after, the chip will have done its job. If the schedule slips a quarter, the demonstration stays an excellent conference paper.
Follow the story on Horizon.


