To make a talking avatar video, upload a clear presenter photo to Morphic, animate the face with subtle motion, add a voice from your script, and lip-sync the words so the avatar appears to speak. Attaching the same photo as a reference keeps the presenter consistent across videos, so you build a repeatable on-screen spokesperson without booking a shoot.
Steps at a glance
- Choose a clear presenter photo
- Animate it as a talking presenter
- Add the script and voice
- Lip-sync the words to the face
- Reuse the avatar and export
Build your talking avatar step by step
1.
Choose a clear presenter photo
Start with a sharp, front-facing photo of the person who will present, well lit and looking toward the camera against a simple background. This one image becomes your presenter, so pick it deliberately and treat it as a casting decision, not a throwaway. Attach it as a reference from the start, because the reference is what holds the face steady across every clip and turns a set of separate videos into one recognisable spokesperson rather than a slightly different person each time.
2.
Animate it as a talking presenter
Animate the photo with the subtle, human motion a presenter has on camera: small head movement, natural blinking, a slight smile, an engaged expression. Keep it understated, because big head turns or exaggerated motion let the face drift from the original. The goal is a presenter who looks alive and attentive while the voice and lip-sync, added next, carry the actual delivery of the message. Overdoing the motion here is the most common mistake, and it is what makes a talking avatar tip from convincing into uncanny.
3.
Add the script and voice
Write what the presenter should say and generate a voice for it, directing the tone and pacing so it sounds like a person presenting rather than a script being read. Pick one voice and keep it for the whole series so the presenter sounds consistent. If you prefer, convert a recording into a different voice with the voice changer, noting that cloning your exact voice from a sample is not yet available.
4.
Lip-sync the words to the face
Combine the animated presenter and the voice with lip-sync so the mouth matches the speech. Clear front-facing footage and clean audio give the most accurate result, so keep the face unobstructed and avoid background noise in the voice track. If the sync drifts, re-run it with a cleaner audio take. The AI lip-sync guide covers the settings in more detail.
5.
Reuse the avatar and export
For a single video, review the resemblance and sync, then export. For a series, reuse the same reference photo and voice on every new script so the presenter stays identical, and assemble longer pieces on the Compose timeline with captions and any b-roll cut in over the presenter. This is what makes an avatar worth the setup: once the photo and voice are dialed in, every future video is just a new script through the same presenter, so a whole training course or content series shares one consistent face. Free exports carry a watermark; a paid plan exports clean and at higher resolution, which matters for training and published content people watch closely.
Talking avatar versus filming a presenter
A talking avatar replaces the shoot, not the message. Here is the trade-off, Morphic first but not the only route.
| AI talking avatar | Film a presenter | Stock presenter footage | |
|---|---|---|---|
| You provide | One presenter photo + a script | Talent, camera, and a studio | A licence and a search |
| Update the script | Regenerate the video | Reshoot with the talent | Not possible; find new footage |
| Consistency | Same face via a reference | Depends on re-booking talent | Different person each clip |
| Best for | Content, training, explainers | High-end brand films | Generic filler shots |
If the presenter is you specifically, how to make an AI video of yourself covers animating your own photo, and adding AI voiceover goes deeper on directing the voice the avatar speaks with.
Put it together
A talking avatar is one photo, subtle motion, a directed voice, and accurate lip-sync, held consistent by a reference. Choose the presenter photo carefully, keep the motion understated so the face holds, and reuse the same reference and voice across every script. Do that and you have a repeatable on-screen presenter you can point at any message, no shoot, no studio, no talent to re-book.