This month has been a whirlwind for AI, with major new models arriving within days of one another. Google just released Gemini 3.8 Flash, and at first glance it has a tough job ahead of it. GPT-6 Astra is OpenAI's most capable model, while Claude Fable 5.1 is Anthropic's latest powerhouse for coding and long-running projects.However, there's one area where Google's new Flash model has an advantage over both: video understanding.Gemini 3.8 Flash can take video directly as an input and reason about what's happening across it. GPT-6 Astra doesn't support video input, and neither does Claude Fable 5.1. And while this might sound like a minor difference, in practice, it actually opens up an entire category of things users can ask Gemini to do that aren't native to the other two models.Gemini can actually watch your video
(Image credit: hocus-focus/Getty Images)Gemini 3.8 Flash accepts text, images, video, audio and PDFs as inputs, with a context window of just over 1 million tokens.Video is particularly interesting because Google isn't simply converting a clip into a handful of screenshots and asking Gemini what's in them.Gemini 3.8 Flash supports what Google calls "agentic video understanding." Instead of processing an entire long video in exactly the same way, the model can navigate through its timeline and decide which transcripts, frames and audio it needs to inspect to answer your question.Google says this approach can use up to 88% fewer tokens on long-form video while delivering roughly 7% higher quality.Sign up to the Tom's AI Guide weekly newsletter summing up all the biggest AI news you need to know. Plus, analysis from our AI editors and tips on how to use the latest AI tools!That means you could give Gemini a long recording and ask something surprisingly specific such as, Where did the speaker mention a particular subject? What happened immediately before someone entered the room? What was being shown on screen when a certain point was discussed?Just from those prompts, Gemini can return answers tied to specific moments in the video rather than forcing you to scrub through it yourself.GPT-6 Astra and Claude Fable 5.1 can't do this nativelyGPT-6 Astra has a slightly larger 1.05-million-token context window and supports text and images, but OpenAI's model documentation explicitly lists both audio and video input as unsupported.Claude Fable 5.1 has a 1-million-token context window and impressive vision capabilities, particularly for understanding diagrams, charts, tables and other visual information. But Anthropic lists its input/output capabilities as "text and images > text."Users can extract frames, generate a transcript or use additional tools before handing that information to the model, it just isn't a native feature in the same way it is with Gemini. Gemini 3.8 Flash was designed to accept it.Audio is part of the advantage, too









