AI

OpenAI’s Jalapeño chip delivers major inference gains in first benchmark results

Close-up of OpenAI's Jalapeño AI inference chip on a circuit board

At the Hot Chips conference on Tuesday, OpenAI shared the first detailed look at its custom inference chip, Jalapeño, along with benchmark results that show significant performance and efficiency gains over current state-of-the-art processors. Tested on SemiAnalysis’s InferenceX benchmark, Jalapeño delivered more tokens per user and higher throughput per kilowatt than Nvidia’s Blackwell system, according to OpenAI’s head of hardware, Richard Ho.

“The bottom line is that the results show a very, very significant performance advance over state of the art,” Ho said during a press call. “Jalapeño can serve more AI work per unit of power, while also returning responses more quickly. It’s very efficient to serve a lot of customers, but it can also be very low latency.”

Also read: General Intuition in talks to raise at $6B valuation, adding Valor and Point72 as it pushes into robotics

The comparison against Nvidia’s Blackwell is notable, but Ho cautioned that the competitive space could shift by the time Jalapeño reaches full deployment. He estimated that Jalapeño would begin shipping in very small volumes at the end of 2026, with more significant deployment following in 2027.

Designed to eliminate inference bottlenecks

First announced in October 2025, Jalapeño was developed in close collaboration with Broadcom, with OpenAI’s own models assisting in the design process. The company plans to treat Jalapeño as a multigenerational platform, allowing AI models, chips, and memory to be developed in concert rather than as separate components.

Also read: The $1.5B Anthropic Ruling Shows AI Copyright Law Is Still Unsettled

That full-stack approach let OpenAI target specific phases of inference that often create delays. In particular, the chip is designed to minimize friction during the prefill and communication phases, which the company says frequently act as bottlenecks in AI processing.

“We designed Jalapeño to minimize data movement and communication delays,” OpenAI said in a blog post presenting the results. “This means that model state, including the KV cache used while generating a response, can be explicitly placed and kept local while the system activates the right combination of compute, memory, and networking for each inference phase.”

By keeping data local and reducing communication overhead, Jalapeño aims to lower latency and improve energy efficiency—two factors that directly affect the cost of serving AI at scale.

What the benchmarks mean for the AI hardware race

The Jalapeño results arrive as major tech companies increasingly design custom silicon to reduce reliance on Nvidia’s dominant GPUs. OpenAI’s move mirrors efforts by Google (with its TPU line), Amazon (Trainium and Inferentia), and Microsoft (Maia), all of which have sought to optimize hardware for their specific AI workloads.

The efficiency gains are particularly relevant for cloud providers and enterprises running large-scale inference workloads, where power consumption and latency directly impact operating costs. If Jalapeño performs as benchmarked, it could give OpenAI a competitive edge in serving its models—and potentially offer those capabilities to external customers through Azure or other cloud partnerships.

However, analysts note that the benchmark comparison is against Nvidia’s current Blackwell systems, and Nvidia is already working on next-generation architectures. By the time Jalapeño scales in 2027, the performance gap could narrow, making the long-term competitive picture less clear.

For now, OpenAI’s hardware push signals a broader industry trend: AI companies are no longer content to rely solely on off-the-shelf processors. Custom silicon allows tighter integration between model design and hardware, potentially unlocking gains that general-purpose chips can’t match.

As deployment ramps up, observers will be watching whether Jalapeño’s real-world performance matches these early benchmarks—and how Nvidia responds with its next-generation offerings.

Neelima Kumar

Written by

Neelima Kumar

Neelima Kumar covers technology and artificial intelligence for StockPil, tracking how emerging tech trends intersect with markets and business.

Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

To Top