Nvidia has moved its specialized inference accelerator, the Groq 3 LPX, into full production. An independent benchmark shows top numbers for token generation, but experts warn the comparison is stacked in Nvidia's favor.

At the Hot Chips 2026 conference, Nvidia announced that its Groq 3 LPX has entered full production. The chip, which Nvidia calls an "interactive AI inference accelerator," extends the Vera Rubin platform and is built to deliver ultrafast token generation for agentic AI systems. Nvidia says it will go live later this year.

In late December, the company paid about $20 billion for the Groq license and brought on founder Jonathan Ross and president Sunny Madra. Groq builds processors tuned for inference rather than AI training.

Speed is important for agentic applications, as agents burn through huge amounts of tokens across hundreds to thousands of inference steps. The faster a system generates tokens, the more reasoning steps and tool calls fit into the same window of time. So within a wait time users find acceptable, an agent can iterate more often, check files, write and test code, and verify results. Nvidia says this cuts coding tasks down to "minutes instead of hours."