AI News

OpenAI publishes first Jalapeño chip benchmarks

OpenAI has published the first measured benchmark results for Jalapeño, its in-house inference chip, claiming better throughput per kilowatt and lower token latency than unnamed commercial systems on a public test.

OpenAI on Monday published the first measured performance results for Jalapeño, its first custom inference chip, in a strategy post titled "The full stack behind abundant intelligence."

On InferenceX, a public benchmark run with GPT-OSS 120B, the company says Jalapeño delivered more peak throughput per kilowatt and lower token latency than "the commercial systems in the comparison." OpenAI also reports strong results on DeepSeek R1 and Kimi K2, which it presents as evidence the gains hold across model families. It describes the chip as a "credible first-party path alongside the accelerators we use from other partners," says co-design spans model, serving software, chip, memory and network, and notes that future generations are already in development.

The post does not name the competing systems, give absolute throughput, latency or efficiency figures, or say how much capacity Jalapeño currently serves or when it scales.

Separately, OpenAI says GPT-5.6 Sol with max reasoning set a new high on the Artificial Analysis Coding Agent Index while using 54% fewer output tokens than "another leading model" — again unnamed.

OpenAI restated its supplier list: Microsoft and NVIDIA as foundational, plus AWS, AMD, Broadcom, Cerebras, CoreWeave, Oracle, SB Energy and SoftBank. It also cited Project Camellia in Georgia, a data center using closed-loop water cooling whose commitments are subject to an annual independent public audit.

The results are first-party and unaudited; independent confirmation on InferenceX would be the next thing to look for.

All stories