A controlled Gemini 3.7 Flash benchmark shows why agentic video is excellent for long-form search—but can cost more than static processing on short clips.

If I only need one number from a long video, why should an AI model sample the entire timeline before answering?

Google launched Agentic Video Understanding on September 1. Instead of processing video at a fixed sampling rate, Gemini can decide whether to inspect the transcript, audio, or selected frame ranges based on the question.

Google reports up to 88% fewer tokens, 66% lower cost, and 7% higher quality on long-form video. Those numbers are compelling, but they do not answer the question I had while building with it:

Is agentic processing cheaper for every video and every query?