ElevenLabs has launched Dubbing v2, an Alpha AI dubbing model designed to preserve a speaker's original performance when audio or video is localized into another language. Rather than depending solely on a transcript, the model conditions directly on the original delivery. The company says this enables translated tracks to retain more of the source speaker's emotion, tone, pacing, emphasis, pitch, and voice identity.

That shift addresses a familiar weakness in automated dubbing. Text can convey the words that were spoken, but it does not fully represent how they were delivered. When translation and synthetic speech are driven principally by text, the result can sound flatter than the source. According to ElevenLabs' official Dubbing v2 announcement, the new model uses the original performance as part of the input so that intonation and emotional intent can carry across languages.

For creators and organizations localizing video, podcasts, campaigns, or other spoken content, the practical goal is more natural multilingual delivery without rebuilding each asset through separate translation, voice, editing, and engineering stages. Dubbing v2 supports more than 90 languages and includes synchronization-aware translation intended to align starts, stops, and pacing with the source material.