Keep one brand character consistent across video, stills, and voice

Build a brand character once and hold it identical across video, stills, and voiceover using reference sheets for the look and a pinned voice for the sound.

A brand character only works if it is the same character everywhere: the same face in a video ad, in a set of stills, and behind the same voice in a voiceover. The hard part is not making the character once; it is making it again, on the next asset, without it quietly turning into someone else. This guide covers how to build a brand character and then hold it steady across all three surfaces, using reference sheets to keep the look constant and a pinned voice to keep the sound constant.

Start with a reference sheet, not a prompt

The mistake that breaks brand consistency is describing the character in words each time. A worded description drifts, because "friendly woman with brown hair" matches a thousand different faces, and the model picks a slightly different one on every run. The fix is a reference sheet: a set of images that define the character, which you then pass into every generation that needs them. The model works from the reference itself rather than paraphrasing it, so the character it produces is recognisably the same one, shot to shot.

This is the mechanism Morphic is built around for consistency, and it is worth investing in up front because everything downstream reuses it.

1.

Generate the character and lock the defining traits

Create the character and settle the details that must never change: the face, the hair, the signature wardrobe or brand colours, the age and build. Be specific about the things that make the character recognisable, because those are what you are about to hold constant everywhere.

2.

Build the sheet with multiple angles and expressions

A single image is a weak reference. Build out a sheet that shows the character from several angles and in a few expressions, so the model has enough of the character to reconstruct it in any new shot rather than guessing at the views it has not seen. This is the difference between a character that holds in a three-quarter turn and one that only works head-on.

3.

Add a location or set reference if the world repeats

If the character always appears in the same environment, a store, an office, a branded backdrop, build a reference for that too. Several references can ride on one generation, so a character, a set, and a wardrobe can all inform a single shot, which is how you keep the whole world consistent, not just the face.

Hold the look across video and stills

Because video and stills draw on the same reference sheet, the character stays the same across both. Generate a still by passing the character reference into an image generation; generate a video shot by passing the same reference into a video generation. The reference is the shared spine, so the person in the product photo is the person in the ad, without you re-describing them each time.

The consistency habits from AI video apply here too. Keep shots short so identity has less time to drift, and when a specific frame is right except for one detail, a targeted edit is usually better than regenerating the whole thing and risking the face. If you find the character shifting between shots, the fix is almost always a stronger reference sheet, more angles, cleaner images, rather than more words in the prompt.

Hold the voice across every voiceover

Visual consistency is only half a brand character. If the voiceover sounds different on every asset, the character still feels inconsistent. Morphic does not clone a specific person's voice, so the approach is to choose one voice and direct it the same way every time. Pick the voice that fits the character, then keep its delivery constant: the same emotion and pacing direction, so the read has the same personality across a launch video, a set of ads, and a help clip.

The reliable way to make "the same, every time" actually happen is to stop re-entering the choices by hand. A workflow step can pin the specific voice and the settings that go with it, so every voiceover you run through it uses the same voice and the same delivery rather than whatever got picked that day. That is what turns a voice you liked once into the character's voice.

Make consistency automatic, not manual

Holding a character steady by remembering to attach the right references and pick the right voice every time works until the day someone forgets. The durable version is to save the process as a workflow: pin the character reference, the model, the aspect ratio, and the voice into the steps, so the consistency travels with the process instead of living in your memory. A teammate can then produce on-brand assets by running your workflow, and the character comes out the same whether it is their first asset or their fiftieth.

Build the character with the AI character generator, settle the voice with the AI voice generator, and save the finished process in the Morphic Workflows library so the whole team ships the same character every time.

FAQ

How do I keep a brand character consistent across different assets?

Build a reference sheet, a set of images that define the character, and pass it into every generation that needs the character, for both stills and video. Working from the reference itself rather than a worded description is what keeps the face, wardrobe, and proportions the same shot to shot. For the voice, choose one voice and direct it the same way each time. Pinning the references and the voice into a workflow makes the whole thing repeatable.

Why does my character look different in every image?

Because a worded description matches many different faces, so the model picks a slightly different one on each run. The fix is a reference sheet instead of a prompt: multiple images showing the character from several angles and expressions, passed into each generation. The more complete the sheet, the better the character holds in views the model has not seen directly, like a three-quarter turn.

Can I use the same character in both video and stills?

Yes. Video and stills draw on the same reference sheet, so passing the same character reference into an image generation and into a video generation produces the same person in both. The reference is the shared spine across the two surfaces, which is what lets the character in a product photo be the character in the ad without any re-describing.

Can Morphic clone a specific person's voice for my character?

No. Morphic does not clone a specific person's voice. The way to keep a character's voice consistent is to choose one voice that fits and direct it the same way every time, keeping the emotion and pacing constant so the read has the same personality across assets. Pinning that voice and its settings into a workflow step is what makes it repeat reliably.

How do I stop the voiceover from sounding different each time?

Fix both the voice and its direction, and stop re-selecting them by hand. Pick the voice that suits the character, set the emotion and pacing you want, then pin those into a workflow step so every voiceover runs through the same voice with the same delivery. Consistency breaks when the choices are re-entered manually each time, so moving them into the process is the durable fix.

How do I make sure my team produces the character consistently?

Save the process as a workflow with the character reference, model, aspect ratio, and voice pinned into the steps. The consistency then travels with the process rather than depending on anyone remembering to attach the right references. A teammate produces on-brand assets by running the workflow, and the character comes out the same whether it is their first asset or their fiftieth.

What is the most common reason brand consistency breaks?

Relying on prompts instead of references. A description drifts a little every run, and those small drifts add up across a campaign until the character no longer looks or sounds like itself. Anchoring the look to a reference sheet and the voice to a pinned choice removes the drift at its source, which is why the setup is worth doing before you produce at volume.