Meta is launching Muse Voice Transcribe, its first real-time audio perception model. It brings multilingual, streaming transcription to Meta AI for Mac, Muse Code, and developers through the Meta Model API.
Muse Voice Transcribe powers system-wide dictation on Mac
Meta Superintelligence Labs says Muse Voice Transcribe combines streaming automatic speech recognition with speaker diarization and endpointing.
In practical terms, it can transcribe speech as it happens, separate speakers across recordings with 20-plus voices, and determine when someone has finished talking, all without a separate post-processing step.
The model was trained across more than 70 languages, with 25 validated at launch. It supports audio longer than an hour, plus native code-switching within or between sentences. Language, keyword, and context biasing can further improve recognition.







