ElevenLabs | Dubbing v2 API Brings Multilingual Localization to Developers

ElevenLabs has made Dubbing v2 available through ElevenAPI, giving developers a way to integrate multilingual audio and video localization directly into their own products and production workflows. The system combines translation, voice preservation, dubbing, and synchronization while focusing on keeping the tone, emotion, and delivery of the original performance across more than 90 languages.


ElevenLabs Dubbing v2 API for multilingual audio and video localization

{getToc} $title={Table of Contents}

Dubbing v2 brings automated localization into creative applications


The Dubbing v2 API is designed to handle more than replacing one voice track with another. ElevenLabs conditions the generated speech on the original performance so characteristics such as emotion, tone, pace, and delivery can carry into the translated version instead of producing a flat or disconnected voiceover.


For creators building tools around video, podcasts, courses, marketing content, or other spoken media, this can reduce the number of separate services needed to prepare localized versions. Translation, voice cloning, dubbing, and synchronization can all become parts of the same automated pipeline.



Voice performance and timing stay closer to the original


One of the important changes in Dubbing v2 is its sync-aware translation system. Translated speech is generated with the timing of the source material in mind, helping the start and end of spoken segments align with the original performance without requiring every section to be manually repositioned.


The model also improves regional accent handling. ElevenLabs highlights distinctions such as Castilian and Latin American Spanish, which can matter when a project needs localization for a specific audience rather than a single generic version of a language. The broader dubbing system currently supports more than 90 languages.


Background audio and multiple speakers remain part of the workflow


Dubbing v2 is built to work with material that contains more than isolated dialogue. It can retain background audio while recreating speech in the target language, reducing the need to rebuild music, ambient sound, or effects simply because the dialogue track has changed.


The system can also detect multiple speakers, including content where voices overlap. For video localization, interviews, collaborative recordings, or narrative content, preserving the separation and characteristics of different speakers is essential if the translated result is going to remain recognizable as the same production.


Developers can automate the pipeline or build more controlled workflows


The API can handle the localization pipeline automatically, which is useful when an application needs to process large volumes of content without requiring manual intervention for every file. Developers can submit audio or video and build the dubbing process into existing publishing, media, or content management systems.


For workflows that require more control, ElevenLabs also supports working with source transcripts and target-language translations before regenerating changed segments. This creates room for human review when terminology, names, brand language, or specific translation choices need to remain consistent across a production.


IMPORTANT: ElevenLabs currently identifies Dubbing v2 as an Alpha model. Creators and developers using it in production should review current API documentation, supported languages, file limits, costs, and model behavior before building a fixed localization workflow around it.{alertWarning}

Daisuki's Take: What This Means for Designers


The most interesting part of the Dubbing v2 API is that localization can become part of a larger creative system instead of a separate final-stage task. A video platform, publishing tool, or internal production workflow can potentially prepare multiple language versions as part of the same process used to manage the original content.


Preserving performance is particularly important for creator-driven media. Translating the words is only part of the job when the personality, pacing, emotion, and recognizable voice of the speaker are also part of the content's identity.


We would still keep human review in workflows where translation accuracy, terminology, cultural context, or precise timing matters. Automation can remove much of the repetitive production work, but the final localized version still needs to communicate the same intention as the original.



Sources and Recommended Links