The benchmark results, presented Tuesday at the Hot Chips conference, mark the first public performance data for OpenAI’s silicon program. Richard Ho, OpenAI’s head of hardware, said the gains represent a meaningful step beyond what current data-center GPUs can deliver for large-scale model serving.
Where Jalapeño Pulls Ahead
The InferenceX benchmark measures real-world inference workloads rather than synthetic compute peaks. Jalapeño registered improvements across two metrics that matter most to AI providers serving millions of queries: tokens delivered per individual user session and total throughput per kilowatt of power consumed.
“The bottom line is that the results show a very, very significant performance advance over state of the art,” Ho said during a press call. “Jalapeño can serve more AI work per unit of power, while also returning responses more quickly. It’s very efficient to serve a lot of customers, but it can also be very low latency.”
That efficiency gap matters because inference — not training — has become the dominant cost for AI companies. As models grow larger and user bases expand, the expense of running queries against trained models now dwarfs the cost of training them. A chip that squeezes more tokens per watt directly improves unit economics for OpenAI’s consumer and enterprise products.
Built With Broadcom, Optimized for the Full Stack
OpenAI first announced Jalapeño in October 2025, developed in close collaboration with Broadcom. The chip is designed as part of a multigenerational platform where AI products, models, silicon, and memory are engineered in concert rather than as separate layers bolted together.
That full-stack integration let OpenAI target specific phases of the inference pipeline that typically create friction. Jalapeño is specifically engineered to minimize delays during the prefill and communication stages — the moments when a model processes an incoming prompt and coordinates data across its network — which OpenAI identified as recurring bottlenecks.
“We designed Jalapeño to minimize data movement and communication delays,” OpenAI said in a blog post accompanying the results. “This means that model state, including the KV cache used while generating a response, can be explicitly placed and kept local while the system activates the right combination of compute, memory, and networking for each inference phase.”
Keeping the KV cache local rather than shuffling it across the network reduces latency at the cost of more complex memory management. OpenAI’s bet is that co-designing the chip alongside its own models makes that tradeoff worthwhile in ways a general-purpose GPU cannot match.
The Timing Problem
Jalapeño’s benchmark victory comes with a significant caveat. The comparison is against Nvidia’s Blackwell platform, which is shipping today. OpenAI does not expect meaningful deployment until the end of 2026, and then only in small volumes. Broader rollout is slated for 2027.
By that point, Nvidia will likely have moved beyond Blackwell. The GPU giant operates on roughly a two-year cadence between major architectures, and competitors including AMD, Google’s TPU team, and Amazon’s Trainium program are all pushing their own inference roadmaps forward. Jalapeño will enter a market that has had 12 to 18 months to respond.
Still, the benchmarks suggest OpenAI’s vertical integration strategy is producing tangible results. If Jalapeño’s second and third generations maintain even a portion of this efficiency lead, OpenAI reduces its dependence on Nvidia for the most expensive part of its infrastructure — a dependency that has constrained margins and supply for the entire AI industry.
What Happens Next
Watch for OpenAI to begin limited Jalapeño deployment in its own data centers by late 2026, likely routing a subset of ChatGPT and API traffic onto the new silicon to validate real-world performance. The 2027 volume ramp will be the real test — both of manufacturing yield with Broadcom and of whether the full-stack co-design advantage holds under production load. Nvidia’s next architecture announcement, expected before Jalapeño ships at scale, will reset the competitive baseline and determine whether OpenAI’s lead survives the gap between benchmark and deployment.
— David Kim, technology desk, AXO News