Google has announced Gemini 3.5 Transcribe, a speech-to-text model that extends beyond dictation in the Gemini app for macOS. The model can use voice commands and screen context to summarize local files, reuse text across apps, and generate images at the cursor. For businesses, the significance is not simply faster transcription. Google is positioning voice as an input layer for work that normally requires moving between documents, applications, and AI tools.
The company announced the model on August 26, 2026. In Google's official Gemini 3.5 Transcribe announcement, it describes the release as its most precise speech-to-text model to date and outlines both end-user workflows and developer access. Gemini 3.5 Transcribe is intended for voice interactions across Google surfaces, including the Gemini app on macOS and Android.
The macOS implementation is the clearest example of the broader product direction. Rather than treating a spoken request as a standalone transcription task, Gemini can combine the spoken instruction with what is visible on screen and call other Gemini models in the background when needed. That enables a workflow such as asking Gemini to analyze a local document, turn selected material into reusable copy in another app, or create an image without manually switching to a separate image-generation interface.






