Audio generation
Coming soon

Seed Audio 1.5

by ByteDance

ByteDance's full‑scene audio model.
Dialogue, music and effects, as separate tracks.

Seed Audio 1.5

Key features

Hear the range

An empty hospital corridor at night with a coat left on a bench
Short drama beatDialogue

Two voices, a corridor hum and a tense pad.

0:00
0:12
A carved stone gate opening in a torch-lit cavern
Game sceneSound design

A warning, dripping water, then the gate opens.

0:00
0:12
Two friends laughing at a kitchen table at night, seen through a rainy window
Podcast openMusic and voice

A short jingle that ducks under the host.

0:00
0:12
A quiet Japanese market lane at dawn with plain paper lanterns
Japanese narrationMultilingual

Japanese narration over a morning market.

0:00
0:12

Technical specifications

6 min

Per generation, up from 2 minutes on Seed Audio 1.0

6

Reference clips per generation, up from 3

30

Including regional variants of Spanish and Portuguese

4

Text alone, or text with voice references, video, or both

Use cases

Two parallel ribbons of amber and blue light travelling side by side through haze

Dubbing and localization

Feed in a finished cut and get a dub that follows the picture, with the original effects and music kept.

A beam of warm light falling across darkness onto a curved panel of glass

Short drama and series

Six minutes covers a full scene in one generation, and separate tracks let you revise a line without new renders.

A stack of thin glass sheets glowing amber at their edges, receding like pages

Audiobooks and audio drama

Narration, character voices and room sound together, with the same cast held across a long chapter.

A bright point of white-gold light with soft rings rippling outward

Ads and promos

Voice-over, music swell and sound effects in one pass. When the client wants changes, regenerate only that layer.

A ring of pale gold light standing upright in dense haze like a portal

Game scenes

Character lines, environment beds and timed effects for cutscenes and trailers, cued to the second.

Two glowing translucent spheres facing each other with light passing between them

Podcasts

An intro jingle, two hosts and a room that sounds real, with each voice on its own track for the edit.

Prompt examples

Two coffee cups and an old ticket stub on a café table by a rainy window

Rainy café reunion

Rainy café, soft piano. Lena, nervous: 'You kept the ticket?' Theo laughs.

0:00
0:12
Edit prompt
A floodlit football stadium at night with the ball in the back of the net

Timed stadium call

[0:00] Crowd builds. [0:03] Commentator: 'He scores!' [0:06] Air horn.

0:00
0:12
Edit prompt
An old wooden door ajar in a dark stone corridor with lightning in a high window

Trailer narration

Low strings, cold wind. Deep narrator: 'Some doors should stay closed.' Thunder.

0:00
0:12
Edit prompt
A hand reaching through a front door for a steaming paper cup left on the doormat

Coffee ad

Pop beat, coffee pours. Bright voice: 'Fresh coffee, at your door by eight.'

0:00
0:12
Edit prompt
A lighthouse on a stony headland at dusk in falling snow

Audiobook opening

Fire crackles, wind outside. Narrator, unhurried: 'One ship was missing.'

0:00
0:12
Edit prompt
Crates of fresh oranges on a sunlit market stall in a town square

Spanish market

Busy market. Vendor, in Spanish: '¡Naranjas frescas, a un euro el kilo!'

0:00
0:12
Edit prompt

Overview

Seed Audio 1.5 is ByteDance's next full-scene audio model and the successor to Seed Audio 1.0. It generates dialogue, music, ambience and sound effects from one prompt, runs up to six minutes, and takes video, voice references and timestamps as direction. ByteDance's partner listings describe it as coming soon.

What Seed Audio 1.5 does differently

The scene comes back in layers. Dialogue, ambience, effects and music arrive as separate tracks, so a wrong line or a loud cue gets fixed on its own. Video input means the sound follows a finished cut, which is what dubbing and translation need.

Seed Audio 1.5 on Morphic

Seed Audio 1.5 is coming to Morphic soon. Seed Audio 1.0 is live now for full scenes, alongside ElevenLabs, Lyria and the rest of the audio models, so you can build the scene on the Canvas today. For prompt structure, see the Seed Audio 1.5 prompt guide, or start from a sample scene in the Seed Audio 1.5 AI audio generator.

Simple pricing

Get started for free today, with the option to upgrade or cancel anytime.

Basic

$19/ month
billed as $0 per year

2400 monthly credits

1 user only

All models

Workflows

Standard

$24/ month
billed as $0 per year

3625 monthly credits

1 user only

All models

Workflows

Pro

$45/ month
billed as $0 per year

6350 shared monthly credits

1 user

+ up to 4 more at extra cost

All models

Workflows

Max

$170/ month
billed as $0 per year

24650 shared monthly credits

1 user

+ up to 9 more at extra cost

All models

Workflows

Enterprise

For higher limits

Custom

pricing and billing terms

High-volume credits
Custom seat limits
All models
Workflows
Pricing Gradient

Free

For playing around

$0

forever free

Up to 20 credits
1 user only
Limited models
Workflows

FAQs

What is Seed Audio 1.5?
Seed Audio 1.5 is ByteDance's next audio model, the successor to Seed Audio 1.0. From one prompt it generates dialogue, music, ambience and sound effects for a whole scene and returns them as separate tracks. It also takes video as input, so it can dub footage or score a cut to match the picture.
When is Seed Audio 1.5 released?
Seed Audio 1.5 is coming soon. ByteDance has shared the feature set with partners but has not opened public access yet. Seed Audio 1.0 is available on Morphic today.
What is new in Seed Audio 1.5 compared with Seed Audio 1.0?
Generations run up to six minutes instead of two, and you can use six voice references instead of three. It adds video as an input, returns dialogue, ambience, effects and music as separate tracks instead of one mix, covers 30 languages instead of 20, and can place sound effects and music cues by timestamp as well as lines.
What are the four input modes in Seed Audio 1.5?
Text only, where the prompt describes everything. Text plus audio, where reference clips set the voices. Text plus video, where the model watches the footage and writes sound to match it. And text plus audio plus video, where picture, prompt and voice references all feed one generation.
Can Seed Audio 1.5 dub or translate a video?
Yes. Give it a video and it generates a dub and soundtrack that follow the visuals across the whole clip. In translation, it replaces the spoken language and keeps the original sound effects and background music. Have a fluent speaker review pronunciation and phrasing before you publish.
Which languages does Seed Audio 1.5 support?
Thirty: Chinese, English, Japanese, Korean, Mexican Spanish, Castilian Spanish, German, French, Brazilian Portuguese, European Portuguese, Thai, Indonesian, Vietnamese, Malay, Filipino, Arabic, Italian, Russian, Dutch, Polish, Turkish, Swedish, Finnish, Danish, Norwegian, Czech, Hungarian, Greek, Romanian and Hindi.
Can Seed Audio 1.5 copy a real person's voice or an existing song?
Not without permission. ByteDance's limits block recognizable imitations of a real person's voice unless you have their consent, along with copyrighted material, brand logos and existing lyrics. You can still cast a voice from your own recordings and write original songs with your own lyrics.
Is Seed Audio 1.5 on Morphic?
Seed Audio 1.5 is coming to Morphic soon. Seed Audio 1.0 is live now, alongside ElevenLabs, Lyria and other audio models. The samples on this page were made on Morphic with ElevenLabs voices and sound effects and Lyria music, mixed into scenes, to show the kind of output Seed Audio 1.5 is built for.