OpenAI is pushing deeper into the hardware side of artificial intelligence with Jalapeño, its first custom-designed inference chip, as the company seeks to make AI responses faster, more efficient and less dependent on off-the-shelf accelerators. Developed with Broadcom, Jalapeño is designed specifically for the inference stage of AI, when trained models generate responses for users.OpenAI says early benchmark results show the processor can deliver substantially more AI work per watt while reducing response latency compared with leading Nvidia systems. The company says the chip is not intended to replace Nvidia across its infrastructure, but to add a specialised layer of computing capacity. Initial deployment is planned for the end of 2026, with later generations already being developed.About The AuthorHey there, i am a technology enthusiast with a deep passion for gadgets, consumer electronics, emerging technologies, and the fast-paced world of digital innovation. Constantly exploring the latest tech trends, product launches, and industry developments, I enjoy translating complex technological advancements into engaging and accessible stories for readers. My interests span smartphones, wearables, artificial intelligence, smart devices, and the broader technology ecosystem. As I begin my journey as a Tech Journalist at Gadgets Now, I am excited to contribute to a platform that informs millions of readers, combining my passion for technology with storytelling to deliver insightful, accurate, and timely tech coverage.The first custom inference chip takes shapeJalapeño is an application-specific integrated circuit, or ASIC, built from the ground up around the requirements of large language model inference. Unlike a general-purpose accelerator, OpenAI says the chip was designed around the kernels, memory movement, networking and serving patterns used by its AI systems.The objective is to generate more useful AI output while consuming less electricity and reducing the time users wait for responses. OpenAI hardware chief Richard Ho said Jalapeño is intended to combine high throughput with low latency, addressing a common trade-off in AI infrastructure. The company developed the processor with Broadcom and says the first design reached manufacturing tape-out in nine months.Optimising the chip around real-world AI workloadsThe significance of Jalapeño lies in its specialised design. OpenAI says its engineers used knowledge from the company's models, software and serving systems to determine how the hardware should handle inference. Inference has two important stages: prefill, where a system processes a user's prompt, and decode, where the model generates an answer token by token.These stages place different demands on computing and memory. OpenAI says Jalapeño was designed to address both, while also improving communication between processors and memory. The company says the architecture keeps important model information, including the KV cache, close to where it is needed, reducing costly data movement.Benchmark gains target speed and energy efficiencyOpenAI's latest results suggest Jalapeño's biggest advantage could be the combination of performance and power efficiency. The company tested the chip against leading Nvidia systems using GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T. OpenAI said Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency across the tested workloads.It also reported between 2.1 and 4.1 times higher performance for highly interactive workloads. These figures are OpenAI's own benchmark results and therefore depend on the testing methodology and comparison systems used.More articles by AuthorTrending StoriesJalapeño’s position against Nvidia BlackwellJalapeño should not be viewed as a straightforward replacement for Nvidia's Blackwell family because the two platforms have different design priorities. Nvidia's Blackwell GPUs are broad AI accelerators capable of supporting both training and inference at large scale, whereas Jalapeño has been purpose-built for inference. OpenAI's published comparisons were against Nvidia GB200 and GB300 systems rather than a direct Blackwell specification-for-specification comparison.That distinction matters when interpreting the results: Jalapeño's claimed advantage is primarily in selected inference workloads, especially latency and energy efficiency, rather than overall AI computing capability.AreaOpenAI JalapeñoNvidia BlackwellPrimary focusLLM inferenceTraining and inferenceDesignCustom ASICGeneral AI GPU platformClaimed inference efficiency1.5–1.9× more work/watt in testsBenchmark baselineClaimed latency1.7–3.6× lower in testsBenchmark baselineStrategySpecialised OpenAI workloadsBroad AI workloadsFaster agents could make the differenceLower inference latency becomes particularly important as AI systems move beyond simple question-and-answer interactions. An AI agent may need to interpret an instruction, create a plan, call a tool, examine the result and then make another decision before completing a task. Each additional step can introduce waiting time.OpenAI says Jalapeño's ability to combine throughput and low latency could therefore make agents more responsive while also improving the economics of serving them at scale. The company argues that better infrastructure could translate into faster ChatGPT responses, more capable Codex tasks and cheaper or more dependable API services as demand for AI continues to grow.The chip is part of a larger hardware strategyOpenAI is presenting Jalapeño as more than a standalone processor. The company says the broader platform combines its accelerator design with memory, networking, software and deployment systems. Broadcom is contributing silicon implementation, networking and connectivity technologies, while Celestica is involved in board, rack and system integration.This full-stack approach gives OpenAI greater control over how its models interact with the underlying infrastructure. OpenAI says the goal is to deploy Jalapeño at gigawatt scale with data-centre partners over multiple generations. The company also says it will continue using Nvidia and other hardware partners, indicating that custom silicon is intended to complement rather than immediately displace commercial accelerators.AI helped accelerate Jalapeño’s developmentAnother notable element is OpenAI's claim that its own AI models helped accelerate the chip's development. According to the company, AI-assisted engineering helped teams explore implementations, optimise parts of the design and shorten development and verification cycles. OpenAI says Jalapeño moved from initial design to tape-out in nine months, a rapid timeline for a high-performance ASIC.The company believes this creates a feedback loop in which AI helps build more efficient hardware, which in turn provides more computing capacity for future AI systems. Jalapeño is the first generation of a planned multi-generation platform, with OpenAI saying its second-generation design is already deep in development. FAQsWhat is Jalapeño and what purpose does it serve in OpenAI's AI infrastructure?Jalapeño is OpenAI's first custom-designed inference chip, created in collaboration with Broadcom. It is specifically designed for the inference stage of AI, aiming to deliver faster and more efficient AI responses while reducing dependency on off-the-shelf accelerators.How does Jalapeño compare to Nvidia systems in terms of performance?Early benchmark results indicate that Jalapeño delivers 1.5 to 1.9 times more AI work per watt and 1.7 to 3.6 times lower end-to-end latency compared to leading Nvidia systems. However, it is not intended to replace Nvidia but rather to complement its infrastructure with a specialized layer for inference workloads.What are the key advantages of Jalapeño for AI workloads?Jalapeño offers significant advantages in terms of both latency and energy efficiency, making it particularly suitable for inference workloads. Its design optimizes communication between processors and memory, which can lead to faster responses and improved performance for interactive AI applications.When is OpenAI planning to deploy the Jalapeño chip?OpenAI plans to begin initial deployment of the Jalapeño chip by the end of 2026, with expectations to scale up deployment during 2027.What role does AI play in the development of the Jalapeño chip?OpenAI claims that its AI models played a crucial role in accelerating the development of the Jalapeño chip by optimizing design implementations and shortening development cycles, demonstrating a feedback loop where AI enhances hardware efficiency.end of article