Google just fundamentally changed how it meters AI usage across its Gemini Apps, and the result is that many users are getting fewer responses than they’re used to. The new system, effective as of May 17, ditches the old fixed prompt-based quotas in favor of a compute-based model that accounts for the actual processing power each request demands.
In English: instead of counting how many times you ask Gemini a question, Google now measures how hard each question makes its servers work. A simple “what’s the weather” costs less than asking the model to analyze a 50-page PDF. The problem is that power users are burning through their allotments significantly faster under this new framework.
How the new system works
The compute-based quotas refresh on a rolling five-hour cycle, but there’s a cumulative weekly hard cap that prevents users from simply waiting out cooldowns and binging on heavy prompts. Usage limits now fluctuate based on prompt complexity, the tools invoked during a session, and overall conversation length.
This change rolled out shortly after Google I/O 2026 and was designed to better reflect the real computational resources being consumed. Users running the advanced Gemini 3.1 Pro model, which is optimized for complex tasks and large data files, found themselves hitting walls much sooner than expected.







