Gemini 3.6 Flash charges you for thinking tokens on top of the answer, and how many it spends is a dial you control per request. On the same 120-word writing task, the default setting billed $0.03316 and the minimal setting billed $0.00110, a 30x swing for output a reader could not tell apart. That dial is the most important cost decision on this model, and it comes with one sharp edge. Gemini 3.6 Flash went generally available on 2026-07-21 at $1.50 per million input tokens and $7.50 per million output, down from $9 output on 3.5 Flash. It shipped alongside Gemini 3.5 Flash-Lite and a security-tuned 3.5 Flash Cyber; this post measures the two general-purpose tiers, 3.6 Flash and Flash-Lite.
TL;DR
reasoning_effort: "minimal" cut per-call cost 91–97% versus the default (a 30x swing on a 120-word task), free on single-step, structured-output, and tool-calling work but breaking multi-step math 3/3 → 0/3.
Google's "17% fewer output tokens" is workload-dependent: our reasoning-heavy tasks ran 19% lighter (32% cheaper), our agent suite 9% heavier (6% cheaper).
The 1M context is real (a needle at 972K tokens recalled) and prompt caching matches Google's published 4,096-token floor exactly, a clean spec match unlike some "1M context" models that undershoot what they advertise.














