Google today introduced Gemini 3.5 Transcribe as its “most precise speech-to-text model yet” that is already powering several first-party products.
This model is “designed to capture your natural speaking style to better understand your intent and recognize custom vocabulary.” As seen in Rambler, Gemini 3.5 Transcribe can remove “ums,” “ahs,” and other filler words from the end result, as well as auto-format your text and allow for natural voice editing. Additionally, Google touts:
More precise transcriptions: Delivers a 4.0% WER (word-error-rate) on streaming audio and 2.6% WER on pre-recorded files across diverse real-world conditions, including background noise and conversational AI interactions, while precisely capturing alphanumeric entities such as postal codes and order IDs.
Custom vocabulary: Recognizes specialized jargon and unique spellings by seamlessly adapting transcriptions to your provided custom vocabulary.
Global language support: Automatically detects and transcribes over 85 languages, seamlessly handling regional accents and diverse dialects.








