OpenAI, in collaboration with Broadcom, is crafting Jalapeño, its inaugural custom inference chip aimed at enhancing AI performance. This innovative processor is designed for rapid and energy-efficient responses, promising impressive improvements in AI processing efficiency and reduced latency for users. Tailored for large language model tasks, Jalapeño is set for its debut by late 2026.

128 chips, 1.7 exaFLOPS, and 27 TB of HBM give Altman and crew a leg up over Blackwell, and maybe even Rubin

Tested on Semianalysis’s InferenceX benchmark, Jalapeño registered both more tokens per user and more throughput per kilowatt than the currently available state-of-the art.