OpenAI's custom Jalapeño chip delivered 1.5x–1.9x more AI work per watt and 1.7x–3.6x lower latency than Nvidia Corp. chips across three models — Gpt-oss, DeepSeek-R1, and Kimi-k2.5-1T-A32B — according to benchmark results the company shared at the Hot Chips conference on Tuesday1,2.
Tested on SemiAnalysis' InferenceX benchmark, Jalapeño registered both more tokens per user and more throughput per kilowatt than what OpenAI described as the currently available state-of-the-art inference processors. The comparison target is an Nvidia Corp. Blackwell system.
"The bottom line is that the results show a very, very significant performance advance over state of the art," said Richard Ho, OpenAI's head of hardware, in a press call. Ho said Jalapeño "can serve more AI work per unit of power, while also returning responses more quickly" and described the chip as offering the "best of both worlds" with lower latency and higher throughput, noting that AI systems typically "have to make a trade-off between the two"3.
Ho estimated that Jalapeño would deploy at the end of 2026 "in very small volumes," with more significant deployment coming in 2027. ANALYSIS The Blackwell comparison carries a built-in caveat: by the time Jalapeño reaches volume deployment, the Nvidia Corp. product line it was benchmarked against will be a generation behind.
Jalapeño is an Application-Specific Integrated Circuit designed for AI inference — the process of running a trained model to complete a task or deploy an agent. The chip was developed by OpenAI in collaboration with Broadcom. OpenAI said its own models assisted in the development process.
The company plans to make Jalapeño a multigenerational platform, allowing AI products, models, chips, and memory to be developed in concert. OpenAI said this full-stack approach enabled the team to address specific phases in the inference process that often cause friction.
Jalapeño was first introduced in June. The Hot Chips presentation marks the first public release of benchmark data for the chip.
The benchmarks land one day after Nvidia Corp. announced the extension of its Vera Rubin NVL72 rack-scale system with fast token generation designed for agentic AI inference and disclosed that its Groq 3 LPX chip is now in full production ctx. ANALYSIS OpenAI's decision to benchmark Jalapeño against Blackwell — rather than the newer Vera Rubin architecture — means the performance claims will face scrutiny as Nvidia Corp. ships its next-generation silicon on a similar timeline.