Microsoft has added MAI-Transcribe-2 to the Microsoft Foundry model catalog in public preview. Announced as part of Build 2026, the speech-to-text model expands Microsoft's first-party MAI lineup across transcription, reasoning, image generation, and voice. For developers and teams evaluating transcription services, the immediate significance is access to an early version of a new Microsoft model through Foundry, rather than a confirmed new benchmark for price, speed, or accuracy.
Microsoft lists MAI-Transcribe-2 among four MAI models entering public preview: MAI-Thinking-1, MAI-Image-2.5, MAI-Voice-2, and the new transcription model. The company's Build 2026 Microsoft Foundry update directs users to try the new models in the Foundry catalog.
That public-preview designation matters. It makes the model available for early testing, but it is not the same as a detailed production specification, general-availability commitment, or published service-level promise. Businesses considering it for recorded meetings, interviews, calls, video captions, or searchable audio archives should treat the release as an opportunity to evaluate fit against their own audio and workflow requirements.
What Microsoft has confirmed about MAI-Transcribe-2













