OpenAI reports first Jalapeño chip results
OpenAI published the first benchmark results for Jalapeño, its in-house inference chip, claiming 1.5–1.9x better performance per watt and up to 3.6x lower latency than unnamed commercial systems on a public SemiAnalysis benchmark.
OpenAI has published the first measured results for Jalapeño, the custom inference chip it announced earlier, testing it on InferenceX, a public benchmark from SemiAnalysis that measures the full process of serving an AI request.
Across three open models — GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T — OpenAI says Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems. For highly interactive workloads it reports 2.1 to 4.1 times higher performance. On Kimi K2.5, the largest model tested, the figures were roughly 1.5x peak performance per watt and 3.4x lower latency. OpenAI says the advantage widened further on its own frontier models in internal testing.
The company frames the design around agentic serving, where many sequential steps compound delay, and says the chip, memory, network and rack-scale system were built together to keep model state — including the KV cache — local and minimize data movement. It credits its own models with helping take Jalapeño from design to tapeout in nine months and with optimizing the chip's arithmetic circuits.
Two caveats sit in OpenAI's own text: the comparison systems are described only as "leading commercially available AI systems" and are not named, and results were normalized using each accelerator's published chip power rating. Jalapeño is rated at 700 watts but drew at or below 550 watts sustained on the tested workloads. OpenAI calls it the start of a multigenerational platform and says it will ramp in the months ahead.