Audio generation

Seed Audio 1.0

by ByteDance

ByteDance's all‑in‑one TTS model.
Voice, music, and sound effects in one generation.

Seed Audio 1.0

Key features

Hear the range

Documentary narrationSpeech

A warm, measured documentary voice-over.

0:00
0:12
Thriller voice-overSpeech

A hushed, tense line read, close and intimate.

0:00
0:12
Spice-market ambienceSound effects

A layered open-air market sound bed.

0:00
0:12
ThunderstormSound effects

A rolling storm building to a distant thunderclap.

0:00
0:12
Orchestral cueMusic

A short rising cue for strings and brass.

0:00
0:12
Lo-fi beatMusic

A relaxed beat with soft keys and vinyl crackle.

0:00
0:12

Technical specifications

ByteDance

Developed by ByteDance's Seed research team.

20

English, Chinese, Japanese, Korean, Spanish, French, German and more.

Up to 2 min

Maximum two minutes of generated audio per pass.

3,000 chars

Room for a full scene brief, dialogue included.

Up to 3

Up to three reference clips, each up to 30 seconds.

Non-streaming

Renders the complete track, not a realtime stream.

Use cases

Audiobooks

Narration, character voices, and sound design for a full book. ByteDance puts the cost near a tenth of studio recording.

Video dubbing

Describe the voice or upload a character image, then use timestamps to land each line exactly where the picture needs it.

Game audio

Character barks, scripted performances, and environmental sound effects for immersive scenes, generated from the script.

One-pass video audio

Give a video clip its narration, sound design, and score in one generation, with no separate mixing step afterward.

Ads and promos

A spoken line, sound effects, and music as one ready-to-use track, made for short-form content.

Dialogue and audio drama

Multiple characters, each with a distinct voice and delivery, in one scene with matching ambience and timing.

Prompt examples

Narrated explainer

Calm narrator, soft kitchen ambience: 'Combine the flour and butter.'

Edit prompt

Timestamp control

Ryan (warm, breathless): '[5.5s:8.0s] Maya! Wait, you're leaving tonight?'

Edit prompt

Audiobook scene

Rain on a library window. Narrator, low and unhurried: 'She read it twice.'

Edit prompt

Sports commentary

Packed stadium, roaring crowd. Commentator, exhilarated: 'OH, HE SCORES!'

Edit prompt

Audio drama

Detective, tense: 'Don't move.' Footsteps stop, a door creaks, a siren.

Edit prompt

Game moment

Deep narrator: 'The ancient seal has broken.' Stone grinding, a dark hum.

Edit prompt

Simple pricing

Get started for free today, with the option to upgrade or cancel anytime.

Basic

$9/ month
billed as $0 per year

900 monthly credits

1 user only

All models

Workflows

Standard

$24/ month
billed as $0 per year

3200 monthly credits

1 user only

All models

Workflows

Pro

$45/ month
billed as $0 per year

6200 shared monthly credits

1 user

+ up to 4 more at extra cost

All models

Workflows

Pro Max

$170/ month
billed as $0 per year

24000 shared monthly credits

1 user

+ up to 9 more at extra cost

All models

Workflows

Enterprise

For higher limits

Custom

pricing and billing terms

High-volume credits
Custom seat limits
All models
Workflows
Pricing Gradient

Free

For playing around

$0

forever free

Up to 20 credits
1 user only
Limited models
Workflows

FAQs

What is Seed Audio 1.0?
Seed Audio 1.0 is ByteDance's all-in-one text-to-speech model. From one text prompt it produces voice, instrumental music, and sound effects together as a finished, mixed track. It works in two headline modes: text prompt to audio (T2A), where everything comes from your description, and text prompt plus audio to audio (TA2A), where you add reference clips to cast specific voices.
What languages does Seed Audio 1.0 support?
Seed Audio 1.0 supports 20 languages: English, Chinese, Japanese, Korean, Mexican Spanish, Castilian Spanish, Indonesian, German, Brazilian Portuguese, French, Thai, Vietnamese, Malay, Filipino, Italian, Russian, Dutch, Polish, Turkish, and Swedish. For best results, write the prompt in the same language as the lines you want spoken.
How does voice reference work in Seed Audio 1.0?
In TA2A mode you supply up to three reference clips of up to 30 seconds each, then tag them in the prompt so each character maps to a recording. The model takes the vocal character and emotion from the reference and carries it across the generation. You can also define a voice from a text description or a character image instead of a recording.
Can Seed Audio 1.0 clone a voice?
Yes. Alongside T2A and TA2A there is a voice cloning mode: upload a single audio clip, and the cloned voice becomes available for straight text-to-speech. ByteDance documents it as a one-clip clone. When the voice has to sit inside a full scene with music, effects, and other characters, use TA2A and its three reference clips instead.
Can Seed Audio 1.0 control the timing of each line?
Yes. Seed Audio 1.0 supports accurate time control: put a timestamp such as [5.5s:8.0s] at the start of a line and the model fits that line to the exact window. This is what makes it usable for dubbing, where dialogue has to match picture.
Can Seed Audio 1.0 generate multiple speakers at once?
Yes. Write a scene with several characters and describe each voice inline. Seed Audio 1.0 gives each speaker a distinct voice, emotion, and pacing in a single generation, along with the ambience and effects around them.
How long can a Seed Audio 1.0 generation be?
Seed Audio 1.0 generates up to two minutes of audio in a single pass, from a prompt of up to 3,000 characters. Longer productions are built by generating scene by scene.
How is Seed Audio 1.0 different from ordinary text-to-speech?
Ordinary text-to-speech picks a voice and reads text aloud. Seed Audio 1.0 goes from text-to-speech to reference-to-audio: one prompt describes the environment, the score, the sound effects, and every character's voice, and the model returns the whole scene mixed together. The difference is scope, a finished audio production versus only the voice.