Changing the voice in a take used to mean recording it again. Speech-to-speech does it differently: you give Morphic a recording you already have, choose a different voice, and get the same performance back in that voice. The words stay the words. The timing stays the timing. The pauses, the emphasis, the rise and fall of the delivery all carry over. Only the voice is new.
This guide walks through the whole flow, from uploading the clip to exporting the result.
What speech-to-speech voice changing is
Speech-to-speech, sometimes called voice conversion, re-voices an existing recording. It reads the performance in your clip and produces that same performance spoken by a different voice. It is not reading a script, and it is not building a copy of anyone. It takes audio in and gives audio out, which is what makes it work on any recording you already have, including the audio inside a video.
It is worth separating this from two things it is often confused with:
- It is not text-to-speech. Text-to-speech starts from a written script and performs it fresh, so the delivery is invented by the model. Speech-to-speech starts from a delivery you already captured and keeps it.
- It is not voice cloning. You are not training a reusable model of a person's voice. You are converting one recording into a target voice for that job. See Is this voice cloning? below for the full distinction.
For a side-by-side on when each fits, see speech-to-speech vs text-to-speech, and the Speech-to-Speech glossary entry for the short definition.
How to change the voice
1.
Open Morphic and start a chat
Go to Morphic and open a new conversation with Copilot. Everything below happens in one place: no separate voice tool, no export-and-re-import between steps. Tell Copilot you want to change the voice in a recording so it knows the job before you hand over the file.
2.
Upload the recording
Attach the clip you want re-voiced. It can be an audio file or a video, since the flow works on the audio inside a video just as well as on a standalone track. Upload the take exactly as you captured it; the delivery in that file is what drives the result.
3.
Describe or pick the target voice
Say what you want the new voice to be. You can describe it in words, a warmer read, an older narrator, a brighter tone, or point to a voice you have already used so the clip matches the rest of a project. The words and the pacing are set by your recording; this step only chooses who appears to be speaking them.
4.
Generate, review, and export
Generate the re-voiced take and listen back. The script, the timing, and the emotion should be the ones from your original clip, now in the new voice. When it lands, export the audio, or if you started from a video, carry the new audio back onto the footage. If it is not quite right, adjust the voice and generate again.
When to reach for it
Speech-to-speech earns its place whenever the performance is already good and only the voice is wrong:
- A take you cannot re-record. The delivery is right, the talent has wrapped, or the moment cannot be staged again. Re-voice it instead of chasing a reshoot.
- A consistent narrator across clips. Several recordings from different sessions or different people can be brought into one voice, so a series sounds like one narrator.
- A different read for a different market or audience. Keep the exact pacing and emphasis of an approved take while changing who is voicing it.
- Placeholder to final. A scratch voiceover recorded on a phone carries the timing; convert it to the finished voice without performing it again.
If you are starting from a script rather than a recording, text-to-speech is the better fit; the comparison page lays out the trade-off.
Speech-to-speech vs the alternatives
| Approach | Starts from | Keeps original delivery | Best when |
|---|---|---|---|
| Speech-to-speech (voice changer) | A recording you already have | Yes, timing and emotion carry over | The performance is right and only the voice needs to change |
| Re-recording | A new session with talent | No, it is a fresh performance | You want a genuinely new read and have the time and talent to capture it |
| Text-to-speech | A written script | No, the model performs it fresh | You have text but no recording, or you need a line that was never spoken |
FAQs
Yes. Speech-to-speech works from your recording, so the words you spoke stay the same words and the timing stays the same timing. Pauses, emphasis, and the overall pacing carry across to the new voice. The only thing that changes is who sounds like they are speaking.
Yes. The flow works on any audio to any audio, and that includes the audio inside a video clip. Upload the video, choose the target voice, and the spoken performance is re-voiced. From there you can bring the new audio back onto the footage.
English and other languages. Speech-to-speech works on multilingual recordings, so a take in a language other than English can be re-voiced too, with the same delivery preserved.
No. Voice cloning builds a reusable model of a specific person's voice that you can then make say anything. Speech-to-speech does not do that. It converts one recording you provide into a target voice for that job; it re-voices a performance rather than creating a standing copy of a person. If you want the distinction in one line: cloning makes a voice you can reuse from text, converting changes the voice in a recording you already have.
Yes. Point each clip at the same target voice and a set of separate recordings comes out sounding like one narrator. If you run this often, you can save the process as a Workflow so the same voice choice is applied every time without setting it up again.
Yes. Emotion lives in the delivery, and speech-to-speech keeps the delivery. A read that was excited, tired, or tender in your recording stays that way in the new voice, because the performance is what drives the conversion rather than a fresh interpretation of a script.
Text-to-speech starts from written words and performs them for the first time, so the model decides the delivery. Speech-to-speech starts from a recording and keeps the delivery you already have. Use text-to-speech when you have a script and no recording; use speech-to-speech when the performance exists and only the voice is wrong. The full comparison covers when each one wins.
No. Uploading the recording, choosing the voice, generating, and exporting all happen in Morphic. If your source is a video, you can assemble the re-voiced audio with the footage on the Compose timeline in the same place, so there is no round trip through a separate editor.