Meta Superintelligence Labs, the division Meta set up to chase artificial general intelligence, has shipped its first product. It’s not a reasoning engine or a world model. It’s a transcription tool.
Muse Voice Transcribe is a real-time speech-to-text model that handles two tasks most transcription services still struggle with: speaker diarization (figuring out who said what) and endpointing (knowing when someone has actually finished talking versus just pausing to think). The model is now available through Meta’s Model API as part of the broader Muse Spark family.
What Muse Voice Transcribe actually does
Real-time transcription sounds simple until you’ve tried to use it in a meeting with more than two people. Most existing tools either mash everyone’s words into a single undifferentiated stream or require manual speaker labeling after the fact. Diarization solves that by automatically attributing speech to individual speakers as it happens.
Endpointing is the other half of the puzzle. It’s the system’s ability to detect when a speaker has genuinely finished a thought, as opposed to taking a breath or collecting themselves mid-sentence. Get this wrong and you end up with transcripts that chop sentences in half or lag behind the conversation by several seconds.








