OpenAI just published benchmark results for Jalapeño — its first custom AI inference chip — and the numbers are credible and significant. Compared to leading NVIDIA hardware (GB200 and GB300), Jalapeño delivers better throughput, better latency, and better power efficiency simultaneously. Not a tradeoff between them. All three at once.
"Jalapeño delivers both higher throughput and lower latency with one architecture, where existing hardware systems often have to make a tradeoff between the two."
That's the claim, and the benchmarks back it up.
What the numbers actually say
Testing ran against the InferenceX public benchmark from SemiAnalysis across three models — GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. None of these are OpenAI-exclusive:










