Audio generation
Available now

Inworld TTS 2

by Inworld

Ninety‑five voices, 100+ languages.
Speech that starts in under a second.

Inworld TTS 2

Key features

Technical specifications

95

Ready-made voices across fourteen native languages

100+

Crosslingual, with one voice identity held across all of them

2,000 chars

Maximum script length per generation

Sub-200ms

Median time to first audio

MP3, 48 kHz

Full-rate audio, ready to drop onto a timeline

2

Inworld TTS 2 and Inworld TTS 2 Flash

Use cases

Two glowing nodes joined by a taut thread of light, one answering the other

Interactive characters

Speech that starts fast enough to hold a conversation. The latency is the feature here, not a footnote.

A long steady band of light with even grain across its whole length

Narration and voice-over

Cast a narrator once and keep the same voice across every cut, pickup, and revision.

A ribbon of light splitting into parallel ribbons that stay in step

Localized voice-over

Run the same script across many languages and keep one voice, so a brand survives translation.

A long slow spiral of amber light receding into depth with even turns

Audiobooks and long-form

Work through a book a passage at a time with the narrator holding steady across chapters.

Many small points of light scattered through depth, joined by faint threads

Game and NPC dialogue

Cast a cast. Ninety-five voices is enough to fill a village without reusing the same one twice.

A clean ascending series of luminous steps receding into depth

Explainers and tutorials

Regenerate a single line when the product changes, instead of rebooking a session.

Prompt examples

A soft bloom of pale gold light opening at the edge of the frame

Warm welcome

The kettle's on. Sit down, tell me everything.

Edit prompt
A gentle fold of bone-white light catching along its crease

Chapter opening

Chapter one. The tide went out that morning and never quite came back.

Edit prompt
A single soft thread of warm ivory light curling once

Recipe step

Combine the flour and the butter until it just comes together.

Edit prompt
Two thin lines of cool white light converging to a point

Breaking news

You are not going to believe who just called.

Edit prompt
A fine mesh of silver-grey light seen very close

Course welcome

Welcome back. Let's pick up where we left off.

Edit prompt
A slow taper of soft amber light narrowing to a point

Tour information

The tour starts at nine, just outside the north gate.

Edit prompt

Simple pricing

Get started for free today, with the option to upgrade or cancel anytime.

Basic

$9/ month
billed as $0 per year

1100 monthly credits

1 user only

All models

Workflows

Standard

$24/ month
billed as $0 per year

3625 monthly credits

1 user only

All models

Workflows

Pro

$45/ month
billed as $0 per year

6350 shared monthly credits

1 user

+ up to 4 more at extra cost

All models

Workflows

Pro Max

$170/ month
billed as $0 per year

24650 shared monthly credits

1 user

+ up to 9 more at extra cost

All models

Workflows

Enterprise

For higher limits

Custom

pricing and billing terms

High-volume credits
Custom seat limits
All models
Workflows
Pricing Gradient

Free

For playing around

$0

forever free

Up to 20 credits
1 user only
Limited models
Workflows

FAQs

What is Inworld TTS 2?
Inworld TTS 2 is Inworld's Realtime TTS-2 text-to-speech model, released on August 31, 2026. It reads a script in one of ninety-five ready-made voices, holds that voice identity across more than 100 languages, and returns audio at real-time speed.
How many voices and languages does it support?
Ninety-five ready-made voices across fourteen native languages, with crosslingual delivery across more than 100 languages. Any voice can speak any supported language, so a voice's native language tells you how it was cast rather than what it is limited to.
Can one voice speak more than one language?
Yes, and it can switch inside a single generation. The voice identity is held across the change, so a bilingual script comes back sounding like one speaker rather than two takes stitched together. That is what makes it practical for localized brand voice-over.
What makes it fast?
Median time to first audio is under 200 milliseconds. That moves text-to-speech out of the render-and-wait category, which is what makes it usable for interactive characters and live agents rather than only for voice-over you generate once.
What is the difference between Inworld TTS 2 and Inworld TTS 2 Flash?
They share the same ninety-five-voice catalog. Inworld TTS 2 gives the fullest read. Inworld TTS 2 Flash is the faster, lower-cost tier for volume work or drafts. You can switch between them without recasting the voice.
How long can a script be?
Up to 2,000 characters per generation. For anything longer, split the script into passages and generate them in sequence. The cast voice stays the same across them, so a long piece stays consistent.
How is it different from ElevenLabs on Morphic?
Both produce human-quality speech on Morphic. ElevenLabs is a full audio suite covering speech, music, and sound effects with fine-grained voice tuning. Inworld TTS 2 is speech only, and its advantages are latency and holding one voice across languages. Many creators use both, one for voice, the other for music and effects.
How do I use Inworld TTS 2 on Morphic?
Open Morphic, switch the prompt bar to Audio, and choose Speech. Pick Inworld TTS 2 as the model, choose a voice, paste your script, then generate. The audio lands on your Canvas alongside your video clips.