OpenAI showed off the first benchmarks for its in-house inference chip at the Hot Chips conference. "Jalapeño" reportedly outperforms both Nvidia's Blackwell and Rubin in throughput per watt and token latency.

The chip handles inference only, meaning it runs AI models but doesn't train them. Jalapeño isn't tuned to OpenAI's own models either. It's a general-purpose LLM inference accelerator.

OpenAI claims Jalapeño delivers 1.5x to 1.9x more AI work per watt at peak throughput across all three tested models, with 1.7x to 3.6x lower end-to-end latency than the best commercially available systems. For interactive workloads, the company says performance is 2.1x to 4.1x higher.

The results come from tests using SemiAnalysis's public InferenceX benchmark. OpenAI provided the numbers. SemiAnalysis verified some runs on-site in the lab. The models tested were GPT-OSS 120B, Deepseek R1 670B, and Kimi K2.5 1T. On GPT-OSS, Jalapeño hit about 1,400 tokens per second per user. On Deepseek R1, it topped 700 tokens per second on a single concurrent request.

At matched decoding speed, Jalapeño achieves 54x to 104x the token throughput per kilowatt compared to the best available accelerator, depending on the model. | Image: OpenAI