Gemini 3.6 Flash vs 3.5 Flash-Lite: API migration and routing guide

Quick answer

Google released gemini-3.6-flash and gemini-3.5-flash-lite as generally available Gemini API models on July 21, 2026. Both support a 1-million-token context window, up to 64K output tokens, thinking, and built-in tools including Computer Use. They solve different production jobs:

Use Gemini 3.6 Flash for coding, multimodal reasoning, and multi-step agent workflows where accepted results matter more than the lowest possible cost.

Use Gemini 3.5 Flash-Lite for extraction, classification, routing, and high-volume subagent work where latency and unit economics dominate.