Google has quietly raised the bar for AI-powered transcription. Unveiled at Google I/O on May 19, 2026, and detailed further with a dedicated model card published on August 26, 2026, Gemini 3.5 Transcribe is the company’s most capable audio processing model to date, and it does considerably more than convert spoken words into text.
The model supports timestamps in MM:SS format, speaker identification, translation, summarization, and emotion detection.
What the model actually does
At 96,000 tokens, Gemini 3.5 Transcribe can process extended audio sessions without losing track of what was said earlier in the recording.
Users can upload common audio file formats, including MP3, directly through the Gemini API, Google AI Studio, or the Gemini macOS application. From there, the model can clean up filler words, identify individual speakers, translate content, or produce a summary, depending on what the user requests.








