How Eleven v4 reads a script
Three choices decide most of the result, in this order: the voice, the words, then the settings.
| Choice | What to do |
|---|---|
| Voice | Pick one whose natural range covers the read. A tag works best when the voice can already do that delivery |
| Words | Write them the way you want them performed. “She said, barely holding it together” or an exclamation mark shifts the delivery without a tag |
| Stability | Lower it for a more expressive read with more variation between takes; raise it for courses where every line should match |
| Resemblance | Controls how closely the output stays to the chosen voice |
| Script length | Up to 3,000 characters per speech generation. Split longer scripts at paragraph breaks |
| Voice type | Prefer a library voice for the final read. Designed voices can sound less polished on v4 |
Eleven v4 audio tags
Audio tags are words in square brackets that shape how a line is performed. They are free text, so treat the table as a starting set and write your own.
| Type | Tags | What it does |
|---|---|---|
| Emotion | [excited], [curious], [sarcastic], [sad] | Colours the whole line |
| Delivery | [whispers], [shouting], [rushed], [slowly] | Changes volume and pace |
| Reactions | [laughs], [sighs], [exhales], [gulps] | Adds a human sound at that point |
| Pauses | [pause], [long pause] | Inserts a gap of that length |
| Sound effects | [applause], [phone buzzing], [light rain] | Plays the sound inside the take |
| Your own words | [low, steady voice, restrained urgency] | Describes the read in plain language |
| Experimental | [sings], [strong French accent] | Works on some voices, test first |
Put a tag right before or after the words it should change, and combine two when a moment needs both, as in [whispers] [sad]. If a tag gets read as a sound, describe the voice instead: [low, gravelly voice] is clearer than [gravelly]. Tags are guidance, so generate two or three takes and keep the best.
Hear the tags






Pacing in Eleven v4 without speed or SSML
Eleven v4 has no speed or style slider and ignores SSML, so timing goes into the script itself. Most v3 scripts carry over once these swaps are made.
| You want | Write this |
|---|---|
| A short beat | An ellipsis... |
| A clear gap (was SSML break) | [pause] or [long pause] |
| A faster or slower line (was the speed slider) | [rushed] or [slowly] before the line |
| More range (was the style slider) | Delivery tags, plus lower stability |
| Emphasis on one word | Capital letters: I said NOW |
| A tricky pronunciation (was SSML phoneme) | IPA between slashes: Niamh /niːv/ |
Eleven v4 dialogue scripts
Choose Dialogue instead of Speech when more than one person talks: two to ten speakers, up to 2,000 characters across the scene. Every speaker hears the whole scene, so replies react to the line before. Tag each turn separately, end a cut-off line with a hyphen, and give speakers clearly different voices.
Maya: [annoyed] The timer went off ten minutes ago. Leo: [defensive] I was just about to- Maya: [pause] [laughs] Is that smoke?
Eleven v4 in other languages
Any voice speaks any of the more than ninety supported languages, and in each one Eleven v4 uses a native accent. Write the script in the target language rather than asking for a translation, and test one line first, since delivery is more fluent in some languages than others.
Eleven v4 vs Eleven v3: what changed
ElevenLabs reports that Eleven v4 took first place on the Artificial Analysis speech arena at launch, and that listeners preferred it about 75% of the time in blind tests (source). For a script, these are the changes that matter.
| Area | What changed in Eleven v4 |
|---|---|
| Delivery | Reads tone and context from the words, so tags fine-tune a performance |
| Consistency | The voice holds steady across retakes, dialogue and long pieces |
| Audio tags | Followed more reliably, including sound effects and your own directions |
| Dialogue | Every speaker hears the whole scene and reacts to the line before |
| Languages | More than 90, up from over 70 on Eleven v3 |
| Accents | A voice speaking a new language uses that language's native accent |
| Controls | Stability and resemblance only. Speed, style and SSML are gone |
Eleven v4 vs Eleven v4 Turbo
Both use the same voices and settings, so a voice cast on one works on the other. On Morphic, pick ElevenLabs for Eleven v4 or ElevenLabs Turbo for the faster tier.
| Eleven v4 | Eleven v4 Turbo | |
|---|---|---|
| Best for | Final narration, ads, audiobooks | Drafts, retakes, high-volume lines |
| Dialogue | Two to ten speakers | Speech only |
| Speed | Full render | Faster; about 150 ms to first speech when streaming, per ElevenLabs |
| Credits | Standard rate | Fewer per character |
Find the delivery on Turbo, then render the keeper on Eleven v4.
FAQs
Write [pause] for a clear gap or [long pause] for a held silence, right where you want it. An ellipsis gives a shorter beat. SSML break tags do not work on Eleven v4, so older scripts need them swapped for tags.
No. Eleven v4 ignores SSML and has no speed or style slider. Use [rushed] or [slowly] for pace, [pause] for gaps, capital letters for emphasis, and the International Phonetic Alphabet between slashes for tricky words.
Yes. Tags are free text, so describe the delivery the way you would brief a voice actor, as in [dry, quietly pleased]. Keep them short, and describe the voice rather than a sound if the model misreads one.
Sound effects, yes: tags such as [applause] or [phone buzzing] play inside the take. Singing is experimental. [sings] works on some voices and not others, so for a full song use a music model on Morphic.
Not by default. Eleven v4 gives a voice a native accent in each language it speaks. An accent tag such as [strong British accent] can steer it back, though results vary by voice.
Usually the voice cannot do that delivery naturally, the tag sits too far from the words it should change, or it was read as a sound. Pick a voice closer to the read, move the tag next to the words, or describe the voice in plain words, then generate a couple of takes.
