Gemini 3.1 Flash TTS
by Google DeepMind
Google's most expressive text‑to‑speech, with audio tags and multi‑speaker dialogue.

Key features
Technical specifications
Multilingual
Style, pace, and accent control across many languages
Up to 2
Two distinct voices in one multi-speaker generation
Audio tags
Natural-language notes plus inline bracket cues
SynthID
Imperceptible AI-provenance watermark on output
Use cases
Video narration and voice-over
Add natural narration to AI or live-action video, with the tone and pacing set in plain language.
Character dialogue
Voice two-speaker scenes for shorts, games, and explainers, each character with its own voice.
Localized voice-over
Narrate the same script across many languages with native pacing and accent control.
Audiobook and long-form
Keep delivery natural and consistent across long passages of narration.
Explainers and tutorials
Clear, directable narration for product walkthroughs, lessons, and how-tos.
Ad reads and promos
Expressive, on-brand voice reads with the energy and emphasis you direct.
Prompt examples
Warm narration
Say this warmly and slowly, like comforting a child: The storm has passed. You're safe now.
Edit promptOverview
Gemini 3.1 Flash TTS is Google's text-to-speech model, announced on April 15, 2026, and integrated into Morphic as an audio model for speech. It turns text into expressive, natural narration that you direct with plain-language instructions and inline audio tags, voices multi-speaker dialogue, and watermarks every clip with SynthID.
What Gemini 3.1 Flash TTS does differently
Most text-to-speech reads your words back at a fixed delivery. Gemini 3.1 Flash TTS takes direction. Write an instruction in plain language before a line to set the tone, pace, or accent, and drop inline cues in square brackets, like [laughs] or [whispering], exactly where you want them. The model performs the direction instead of reading it aloud, so a single script can shift from a hushed aside to a full-energy read.
Gemini 3.1 Flash TTS on Morphic
On Morphic, Gemini 3.1 Flash TTS sits in the audio model picker for Speech, alongside ElevenLabs. Switch the prompt bar to Audio, choose Speech, pick Gemini 3.1 Flash TTS, then write your script with any direction or tags, choose a voice and language, and generate. The audio drops straight into Canvas next to your video clips.
Simple pricing
Get started for free today, with the option to upgrade or cancel anytime.
Basic
2400 monthly credits
1 user only
All models
Workflows
Standard
3625 monthly credits
1 user only
All models
Workflows
Pro
6350 shared monthly credits
1 user
All models
Workflows
Max
24650 shared monthly credits
1 user
All models
Workflows
Enterprise
For higher limits
Custom
pricing and billing terms

Free
For playing around
$0
forever free
FAQs
Flux 3 Image
Black Forest Labs
Black Forest Labs' control-first image model. Place every element, edit one box at a time.
Ideogram 4.5
Ideogram
Ideogram's precision edit model. Stack edit on edit, with no drift or artifacts.
Eleven v4
ElevenLabs
ElevenLabs' most expressive voice model. Audio tags, 90+ languages, and a faster Turbo tier.
Kling 4.0 Flash
Kling
Kuaishou's speed-tuned Kling 4.0 for high-volume video. Fast 3 to 20 second clips at 720p.