Google's new Gemini 3.5 Transcribe recognizes over 85 languages, strips filler words, and corrects slips of the tongue in real time. It hits a 4.0 percent word error rate in streaming mode, with 70 percent lower latency than its predecessor, Chirp 3. Through function calling, the model can hand off tasks to other Gemini models.

Now you can get more intelligent speech-to-text transcription with Gemini 3.5 Transcribe.

Google today introduced Gemini 3.5 Transcribe as its “most precise speech-to-text model yet” that is already powering...