(Image credit: Nvidia)
This Tom's Hardware Premium article is free to read with a Tom's Hardware account; no payment necessary. We're offering free access from August 23 to 26 so you can read all of our reporting from Hot Chips.Groq's former chief architect stood on stage at Hot Chips 2026 and presented his former company's inference chip as Nvidia silicon. Igor Arsovski, now Nvidia's VP of hardware, presented the Groq 3 LPX rack's architecture and published the first third-party benchmark of the hardware: Artificial Analysis measured it at 3,431 output tokens per second on a 100K-context Gemma 4 31B reasoning workload, roughly four times the 870 tokens per second of the next-fastest public endpoint. Arsovski said the rack is already in production, built on the LP30 chip Nvidia obtained through its $20 billion Groq deal in December 2025, the same deal that pushed the Rubin CPX it replaced off Nvidia's roadmap.SRAM without HBMArtificial Analysis ran the comparison on a private, pre-release Gemma 4 31B endpoint served through Google Cloud, taking the median of 50 sequential client requests at a concurrency of one, while the public providers it measured against ran shared production serverless endpoints. Serving one request at a time produces the highest per-user token rate the hardware can post, and it's not directly comparable to the multi-tenant conditions the other endpoints run under.Nvidia's on-stage demo showed a higher figure still, 10,996 tokens per second on the same 31B model, which Igor Arsovski, Nvidia's VP of hardware, flagged on stage as "self-reported" before telling the audience the aim was "third-party verified independent benchmarks that you guys can trust." Gemma 4 31B is also a dense model small enough to sit inside a single LPX rack, and the picture at trillion-parameter mixture-of-experts scale, where memory capacity becomes the main constraint, went unaddressed.










