Google DeepMind’s Gemma team and Hugging Face ran an experiment they called the “Fast Gemma Challenge,” and the results are worth paying attention to: a 5x improvement in inference speed for the Gemma 4 model, achieved by more than 100 AI agents and human participants working together over roughly six days.
The peak performance hit 491.8 tokens per second.
How the challenge actually worked
The Fast Gemma Challenge ran from June 26 to July 2, with the specific target being the Gemma 4 E4B-it model. The constraint was deliberately tight: participants had to optimize inference using a single NVIDIA A10G GPU with just 24 GB of memory.
Participants submitted their optimizations through an OpenAI-compatible endpoint, and progress was tracked via a public leaderboard and shared dashboard.






