Dubbing takes one recorded performance and makes it land in a new market. Localization is the wider job around it: the captions, the timing, the voice, and the sync that make a video feel like it was made for the audience watching it, not translated at them. Neither is a single button. Both are a short sequence of steps, and on Morphic those steps live in one place, so a clip does not leave and come back between each one.
This guide walks the workflow end to end using only what Morphic actually does: transcribe the source, translate the captions, re-voice the audio, sync it to the footage, and assemble the cut.
What AI dubbing and localization involve
Dubbing replaces the spoken audio of a video with a new-language or new-voice version that matches the picture. Localization surrounds that with the transcript, the translated captions, and the final assembly. The parts that used to mean separate vendors and separate tools, transcription, translation, voice, lip-sync, and the timeline, are steps you run in sequence rather than files you ship around.
Two honest notes before you start:
- This is a workflow, not one-click dubbing. You move through the steps and check the result at each one. That is a feature: you catch a mistranslation or an off read while it is cheap to fix, not after everything is rendered.
- The original voice is not cloned. When you re-voice with speech-to-speech, you are converting the recording into a target voice, not building a standing copy of the original speaker. See the FAQ for the distinction.
The workflow, step by step
1.
Transcribe the source to an SRT
Start by turning the spoken audio into text. Upload the video and have Copilot transcribe it; the spoken words become a downloadable SRT with timing, and the language auto-detects. This transcript is the spine of everything that follows, because every later step works from these words and these timings.
2.
Translate the captions for the new market
Translate the transcript into the target language. You get captions in the new market's language, timed to the original, which you can review and correct before anything is voiced. Getting the text right here is what keeps the rest of the job honest: the voice and the sync are only as good as the words they are built on.
3.
Re-voice the audio
Now produce the new spoken audio, one of two ways depending on what you want to preserve:
- Speech-to-speech keeps the original performance. It re-voices the existing recording into a different voice while holding the delivery, pacing, and emotion of the take. Reach for this when the original read is the thing worth keeping.
- Text-to-speech performs the translated script fresh. Reach for this when you are moving into a new language and want the line spoken naturally in that language from the translated text.
The speech-to-speech vs text-to-speech comparison lays out which to pick for a given job.
4.
Lip-sync the new audio to the footage
With the new audio in hand, sync it to the picture so the mouth matches the words. Lip-sync aligns the speaker's mouth movements to the re-voiced track, which is what stops a dub from feeling like a dub. This is the step that makes the localized version read as native rather than overlaid.
5.
Assemble on the Compose timeline
Bring it together on Compose: the footage, the new audio, and the captions if you are burning them in. Set per-clip volume so the voice sits right against any music or effects, add transitions where the cut needs them, and export. The finished piece leaves in the format the market needs, from the same place you started.
Why one workspace matters here
The reason dubbing has historically been slow is not any single step; it is the handoffs between them. A transcript goes to a translator, the translation goes to a voice vendor, the voice comes back and goes to an editor, and each handoff is a place to lose time and context. Running transcription, translation, voice, lip-sync, and the timeline in one workspace removes the handoffs. Nothing is exported, re-uploaded, or re-explained between the steps, which is what turns a multi-vendor job into an afternoon.
Speech-to-speech vs text-to-speech for dubbing
| Speech-to-speech | Text-to-speech | |
|---|---|---|
| Starts from | The original recording | The translated script |
| Keeps the original delivery | Yes, pacing and emotion carry over | No, the model performs it fresh |
| Best for | Keeping a strong original performance in a new voice | Moving into a new language from translated text |
FAQs
If you re-voice with speech-to-speech, yes. Speech-to-speech works from the original recording, so the delivery, pacing, and emotion of the take carry into the new voice. If you use text-to-speech from the translated script instead, the read is performed fresh, so the emotion is set by how you direct that read rather than carried from the source. See speech-to-speech for the mechanic.
Transcription auto-detects the source language, captions can be translated into another language for the target market, and the re-voicing works in English and other languages. That covers taking a source clip in one language and delivering it spoken and captioned in another.
No. Transcription, caption translation, re-voicing, lip-sync, and the Compose timeline are all in Morphic. The point of running them in one place is that the clip does not leave and come back between steps, so there are no handoffs to a separate transcription service, translator, voice vendor, or editor.
No, and that is deliberate. Dubbing is a sequence of steps, transcribe, translate, re-voice, sync, assemble, and you review the result at each one. That lets you catch a mistranslation or an off read while it is cheap to fix, rather than discovering it after the whole video is rendered.
Use speech-to-speech when the original performance is worth keeping and you want that same delivery in a new voice. Use text-to-speech when you are moving into a new language and want the translated line spoken naturally from text. Many localizations use both across a project. The full comparison covers the trade-off in detail.
No. Re-voicing with speech-to-speech converts the recording into a target voice for that job; it does not build a reusable copy of the original speaker. You are changing the voice in a recording, not training a model of a person's voice.
Yes. The translated captions can be styled and burned onto the clip, or exported as an SRT to ship alongside it. If you burn them in, you do it on the same timeline where you assemble the final cut, so captions, audio, and footage all come together in one export.
Yes. Once a dubbing sequence works, you can save it as a Workflow and run the same steps on the next clip, so a localization process you settled on once becomes something you run rather than rebuild each time.