Seedance 2.5: complete guide, features, prompts, and longer video

Seedance 2.5: complete guide, features, prompts, and longer video

The complete Seedance 2.5 guide: features, longer single-take video, reference-led continuity, camera control, native audio, and prompt examples.

Seedance 2.5 features and capabilities

Seedance 2.5 is ByteDance's next-generation video model. It extends the Seedance 2.0 line with longer native takes, many more references, and native audio across 10+ languages, while keeping the AI camera control the line is known for.

Seedance 2.5 is rolling out, so some details may change before or at its full launch, which ByteDance expects on August 7, 2026.

FeatureWhat it doesBest for
Longer single-take videoHolds one continuous shot far past a few-second beat, up to 30 secondsAd spots, short scenes, long reveals
Reference-led continuityTakes up to 50 multimodal references to lock a character, set, and paletteSeries, multi-shot sequences, brand work
In-clip editingChanges one element of an existing clip without a full re-renderFixing a take, targeted revisions
AI camera control with audioDirects the camera in plain language, with native audio in the same passDirected shots, sound-on scenes
Audio-only referenceA voice, music, or sound-effect track can drive pacing, beat-matching, and lip-syncMusic videos, lip-synced dialogue, localized voiceover

Longer single-take video

The headline change is length. Seedance 2.5 holds a single continuous take far past the few-second clips most models produce, so a full reveal or a short scene can play out in one shot instead of several stitched together. You write the prompt as one evolving motion, describing how the subject and camera move across the whole take rather than a single frozen frame.

Reference-led continuity

Seedance 2.5 leans on multimodal references to keep a look steady. Supply an image of a character, product, or set, and the model carries that look through the clip and from shot to shot. It accepts up to 50 reference inputs per generation (30 images, 10 video clips, and 10 audio), spanning images, video clips, audio, and even simple 3D white models, so a team can lock a character, a location, and a palette together, which is what keeps a sequence recognizable rather than drifting scene to scene.

In-clip editing and extension

Editing is targeted rather than a full re-roll. For a single change to an existing clip, adjust that element in place and the rest of the frame stays put, which is faster than regenerating the whole shot and avoids losing a take you already like. You can point at a moment by timestamp, so a change lands only where you want it.

OperationWhat it doesExample
Instruction editAdds, removes, or changes content by text, optionally at a set timestampReplace the scene from 0:02 to 0:05, keep the motion and rhythm
Reference-image editSwaps an element to match a supplied imageReplace the jacket with the one in the reference, keep the rest
Add a subjectDrops a person, prop, or effect into the frame, rest unchangedAdd an energy bow into the character's hands
Remove a subjectErases an object, watermark, or logo, then fills the gapRemove the drone and its track, inpaint the sky
Audio editReplaces the music, adds effects, or changes the voiceSwap the background music, keep the visuals
Extend a clipContinues forward, fills the moment before, or bridges two clipsExtend six seconds forward as the sky darkens

Extension is high-fidelity: it can continue a clip past its last frame (forward), generate the moment before its first frame (backward), or build a transition that connects two clips, keeping character, rhythm, and style intact. You set how many seconds to add.

AI camera control with audio

Camera moves are directed in plain language, a slow push-in, a low tracking shot, a pull-back, so motion reads as an intentional camera rather than drift. Native audio generates in the same pass, so a scene can come back with room tone or effects already in place instead of a silent clip.

Native audio and 10+ languages

Sound generates in the same pass, so a scene can come back with voice, music, or effects already in place. New in 2.5, an audio-only reference lets a single voice, music, or sound-effect track drive the pacing, beat-matching, and lip-sync of the shot. Seedance 2.5 also generates natively in 10+ languages and follows instructions more closely, so you can describe a scene, and localize its voiceover and titles, in your own language.

Seedance 2.5 use cases

Long single-take establishing shots

One unbroken move across a landscape, held far past a few-second beat. A long native take carries a whole reveal without a cut, so an aerial reads as a single continuous camera rather than stitched clips.

Detailed close-up craft

Hands at work, jewelry, and mechanisms hold their fine detail through the move. Reference-led generation keeps the object consistent frame to frame, so a close inspection shot stays crisp instead of smearing.

Sports and athletic motion

Fast bodies and shifting weight stay coherent across the clip, with motion that tracks the action rather than blurring it. A dawn sprint reads with real timing and follow-through.

Cinematic sci-fi scenes

Big sets, atmosphere, and volumetric light give a shot scale. Camera control moves through the space with intent, so a hangar interior feels staged for a scene rather than a static render.

Travel and documentary

Wide vistas and slow reveals finish clean at high resolution, ready for a large screen. A desert crossing holds its detail from foreground grain to the far horizon.

Surreal, imaginative concepts

Impossible scenes hold together because the subject stays consistent through the shot. A whale drifting past a diver keeps its scale and weight, so the concept lands instead of falling apart mid-move.

How to get the best out of Seedance 2.5

Seedance 2.5 rewards a clear shot brief and a habit of leaning on references for continuity. A few practices carry most of the quality:

  • Write for one continuous take. Describe how the subject and camera evolve across the whole length of the shot, not a single instant.
  • Name the camera move. A specific "slow push-in" or "low tracking shot" reads as intent, where "cinematic" tells the model nothing.
  • Lock the subject with a reference. Supply a clear image of a character, product, or set so the look carries through the clip instead of drifting.
  • Say what each reference is for. Name the character, the set, and the palette so the model knows which reference drives which part of the scene.
  • Reuse references across shots. Carrying one set of references from cut to cut is what keeps a sequence consistent.
  • Stage complex blocking with a 3D whitebox. A simple white-model scene passed as a reference lets you lay out the set and the camera path before you spend a render on it.
  • Edit in place, don't re-roll. For a targeted fix to an existing clip, change that element rather than regenerating the whole shot.
  • Drive rhythm with an audio-only reference. Attach a music or voice track and the model can match the pacing, beats, and lip-sync to it.

Choosing references for stable results

More references is not always better; a focused set is what holds quality:

  • With image references, keep the main subjects to 8 or fewer, and push to 9–12 only when you have to.
  • With a video or audio reference, keep subjects to 5 or fewer, and aim for a 5 to 10 second clip: long enough to read, short enough to stay clean. Video references total 30 seconds across all of them, and so does audio.
  • When editing an existing video, keep the source clip under 20 seconds and pair it with 1 to 5 reference images, pushing to 6–8 only when the edit needs them.
  • With five or fewer subjects, single or multiple views both work; past five, prefer single-view images, or upload separate angles rather than one combined image. Separate angle images hold up better than several views collaged into one picture.
  • Prioritize in this order: core characters, then key products or props, then the scene, then overall style.

For the full capability list and specifications, see the Seedance 2.5 model page.

Seedance 2.5 prompt guide

A strong video prompt reads like a short shot brief, not a caption, so the model has a subject, a motion, and a camera to work with rather than a still frame in words. ByteDance's own prompt guide sets out one formula for this, and everything else on this page builds on it.

The Seedance 2.5 prompt formula

Six elements, in this order, and you drop any you do not need:

ElementWhat it coversRequired
Subject and actionWho or what is in frame, and what they doYes, this is the foundation
Scene and environmentLocation, time, weather, spatial relationships, backgroundOptional
Visual styleLighting, color, materials, texture, overall moodOptional
Camera movement or cutsShot size, angle, movement, focus subject, transitionsOptional
AudioDialogue, voice, ambience, sound effects, musicOptional

Written out, that is four short lines:

[Subject] performs [primary action or event] in [scene and environment].
The visuals feature [visual style].
Use [shot size, camera angle, camera movement, or cuts].
Audio includes [dialogue, ambience, sound effects, or music].

A worked example:

A ceramic artist finishes a pale blue cup in a studio at dawn, lifts it from
the wheel, and places it in the center of a wooden shelf.
Soft morning light enters through the window. The wet clay has a delicate
sheen, and the workbench remains tidy.
Begin with a medium shot of the wheel-throwing process, slowly push in toward
the cup's surface texture, then cut to a frontal view of the shelf.
Retain the low hum of the pottery wheel, the friction of clay, and subtle
indoor ambience.

Aspect ratio, duration, and the other generation settings stay out of the prompt. You set those on the generation page, not in the text.

Audio, dialogue, and subtitle syntax

Plain language works, but four bracket types tell Seedance 2.5 exactly which kind of sound or text you mean. They are the cheapest way to stop a line of dialogue being drawn on screen as a caption, or a sound effect being spoken aloud.

ContentBracketExample
Music( )(Soft, rhythmic piano music plays in the background)
Sound effect< ><A bell rings in the distance>
Dialogue{ }{Hello, welcome back.}
Subtitle【 】【Chapter One: Departure】

Name the language before any line that is not in Chinese, and name it again if the model keeps reaching for the wrong one. The pattern is language, then regional variety or accent, then delivery style, then speaker, then the line:

Dialogue language: American English. The girl says in natural, conversational
American English: {I thought you weren't coming.}

That last part matters more than it looks. Naming a regional variety, "authentic Los Angeles English" rather than just "English", is what gets you a specific delivery instead of a neutral read.

Naming references and defining their roles

References are where Seedance 2.5 pulls away from earlier models, and where most weak prompts fall down. Uploading an image is only half the instruction. The prompt has to say what that image contributes, and what to ignore in it:

@Image 1 defines the ceramic artist's facial features, hairstyle, and dark
green apron. Do not use the image background.
@Image 2 defines the wooden workbench, window placement, and morning light of
the pottery studio. Do not use the people in the image.
@Video 1 defines the pacing of throwing clay with both hands, lifting the cup,
and placing it down. Do not use the person's identity, clothing, or scene.

Three rules carry most of the reliability here:

  • Bind one reference to one subject at a time. Write out each mapping. A line like "@Images 1 through 4 define four characters respectively" never says which image is which character.
  • Add the exclusions. Backgrounds, bystanders, and compositions ride along from a reference unless you rule them out.
  • Say when several images are one thing. Four angles of the same lamp need "all four images define one folding desk lamp, and the output must contain only one lamp throughout", or you get four lamps.

When a reference video already carries the motion and camera work, name only the attributes to inherit. Restating the action in words fights the reference you attached.

Directing many references at once

Past a handful of references, the goal stops being description and becomes casting. Work in this order: define each material's role, map subjects, group by type, build profiles for the important ones, then pick references per scene.

Grouping keeps a long list readable:

[Characters]
<Conservator> corresponds to @Image 1. Use only the appearance,
hairstyle, and clothing.
<Registrar> corresponds to @Image 2. Use only the appearance,
hairstyle, and clothing.
Do not interchange their appearances, clothing, actions, or dialogue.

[Props]
<Sample Case> corresponds to @Image 5 and belongs only to <Conservator>.

[Scenes]
<Conservation Lab> references @Image 7. Use only the space,
materials, and lighting.

For a character who recurs across scenes, write one profile that gathers every reference attached to them, including a "do not use" line. Then select per scene, so each shot draws only the references it needs:

Scene 1 | Inspection in the Conservation Lab
Use: <Conservator>, <Sample Case>, <Conservation Lab>, and the
case-opening motion from @Video 1.
Event: <Conservator> opens <Sample Case> at the workbench and
inspects the sample inside.
End state: <Conservator> remains on the inner side of the workbench,
<Sample Case> beside their right hand.

Fifty references is a casting pool, not a shopping list. The point is helping the model pick the right materials for the shot in front of it, not getting everything on screen at once.

Longer, edited, and extended Seedance 2.5 shots

Length is what Seedance 2.5 is built for, and a 30-second take needs a different prompt shape than a five-second one. So do the edit and extension modes, which run on the clip you already have rather than a blank frame.

Writing a 30-second Seedance 2.5 video

Split the story into consecutive stages. Give each stage exactly one main change, and state what should be visible on screen when it ends. The end state is what carries continuity from one stage to the next:

[Generation Goal]
Generate an instructional video showing a flower shop's order-packing process.

[Stage 1]
Initial state: <Florist> stands behind the workbench. Loose stems,
scissors, and wrapping paper lie on the tabletop.
Primary event: <Florist> arranges the stems and trims them to length.
End state: <Florist> holds the bouquet in the left hand, scissors
back on the right of the workbench.

[Stage 2]
Continue from the previous stage: both characters keep the same
identities and clothing.
Primary event: <Store Assistant> unfolds the wrapping paper,
<Florist> places the bouquet inside and ties it.
End state: the wrapped bouquet lies flat in the center of the
workbench, bow facing camera.

[Maintain Consistency]
Keep character identities, clothing, prop ownership, and workbench
orientation consistent.

Timing and pacing

Stages are the default. Reach for timestamps only when a handoff, entrance, transition, or beat has to land at a particular moment.

PatternUse it forExample
Time rangeBudgeting pacing across the clip0–3 seconds… 3–7 seconds… 7–12 seconds…
Exact time pointOne critical beatAt 5 seconds, the camera whip-pans left and completes the transition
Relative timingA delay between two eventsThree seconds after the character presses the button, the lights go out

Keep ranges consecutive and non-overlapping. They are time budgets rather than edit points, so an action can land slightly either side of a boundary. Too little content in a range gives the model room to wander; too much causes rushed cuts or dropped events. Asking for three actions in one second does not work.

Editing an existing clip

An edit is a patch, not a new prompt. Name the source clip as the sole master, state the scope of the change, and list what must not move:

[Edit Goal]
Edit @Video 1. Only from 4-7 seconds, change the cool blue light on
the right wall to warm orange light.

[Source Video Role]
@Video 1 is the sole editing master. It defines the character, room
layout, actions, composition, camera movement, audio, and event order.

[Edit Scope]
Change only the light color on the right wall and the area it
illuminates. Allow skin tone to respond naturally.

[Content to Preserve]
Keep the character's identity, clothing, expression, position, motion,
room structure, camera movement, dialogue, and ambience from @Video 1.

Swapping a subject adds one more block, and it is the block people forget. The replacement has to inherit the original's timeline: every appearance, movement, occlusion, and exit, with the same timing, path, and speed. Without it you get the new object in the right frame and the wrong rhythm. Background replacement works the same way, scoped to everything outside the subject's silhouette.

Audio edits are separate from visual ones. Name the speaker or sound category, the change, and what stays: removing the music while keeping dialogue, lip sync, ambience, and effects is a single instruction.

Extending a clip forward or backward

Extension bolts new footage onto a boundary frame, so describe that frame before you describe the new action. Forward, the extension's first frame continues from the source's last frame. Backward, the extension's last frame has to arrive at the source's first frame:

@Video 1 is the source video to extend forward.

Extend @Video 1 forward. The first frame of the extended segment directly
continues from the last frame of @Video 1. Maintain the same locked-off medium
shot, the paper airplane's position and orientation, the classroom-window
background, the afternoon lighting, and its movement toward frame right.

Then, the paper airplane continues gliding right and exits the frame while the
curtain beside the window sways slightly.

Backward extensions need one extra guard. Say which materials belong only to the original clip, or characters and props that should arrive later turn up early in the segment you just generated. And treat "connect to the source video" as insufficient on its own: spell out the source's first frame as the explicit end state, or the image keeps changing after it has already arrived.

Locked settings you cannot override

Editing, extension, and first-frame generation each lock some parameters to the material you fed them.

TaskAspect ratioDuration
Video editingInherited from the input video, cannot be setInherited, cannot be set, within about 0.3 seconds
First or first-and-last frameInherited from the first imageCan be set
Video extensionInherited from the input video, cannot be setCan be set

The practical catch is on first-and-last-frame work: give both images the same aspect ratio, or the last frame gets stretched to match the first.

Advanced Seedance 2.5 direction

Keyframes, storyboards, and 3D blockouts

You can hand Seedance 2.5 the structure of a shot rather than describing it. Four ways, in rising order of control:

  • First and last frames. Say in the first line that one image is the first frame and another is the last, then describe the single continuous action between them. Describe each anchor separately; a combined "these two are the first and last frames" does not bind either one.
  • Multi-keyframe sequences. Open with "use @Image 1 through @Image N as keyframes in this order", then give the key state each image represents. Separate images align better than several frames collaged into one grid, and they control stage order rather than reproducing every frame.
  • Storyboard grids. A grid communicates shot order and rough composition, not exact detail. Keep it under about 15 panels, use clean line art, minimize text labels, state the reading order, and rule out the grid's own style so the line art does not end up in the render.
  • 3D blockouts. A coarse blockout carries timing: paths, blocking, camera movement, cut points, lighting changes, and sound rhythm. Map each grey shape to its final subject by name. A fine blockout already has complete structures, so it is for re-rendering materials, characters, scenes, and style while the structure and camera work hold. Strip path lines, coordinate axes, and camera frustums out of a fine blockout first, or they get rendered too.

Directing performance and emotion

Emotion words set a direction but leave the acting open. What tightens it is naming what a viewer would actually see or hear: eye movement, brow tension, mouth movement, breathing, gaze direction, hand movement. Two to four cues is usually enough for a single emotional turn, and you never need to list every facial detail.

The default shape is a trigger, an immediate reaction, a gradual change, then the settled expression:

The overall emotion shifts from [starting emotion] to [ending emotion].
After [triggering event], [subject] first shows [immediate observable reaction].
Then, [eyes, brows, mouth, breathing, gaze, or hand movement] gradually [changes].
Finally, [subject] expresses [target emotion] through [outward behavior].

For an emotion that turns several times, chain it to events instead: each new thing the character hears or sees moves the performance one step.

Camera language Seedance 2.5 understands

Standard vocabulary can go straight into the prompt: shot sizes from extreme wide to extreme close-up; push in, pull out, pan, lateral move, follow, orbit, dive, dolly out, tilt up, handheld shake; low angle, overhead, first person.

Popular techniques work too, as long as you say which subject they apply to and where the move starts and ends. A one-take shot needs the spaces it passes through in order; a dolly zoom needs the subject size to hold and which way the background should travel; bullet time needs the action to freeze and the orbit direction; a speed ramp needs where the action accelerates and where it settles.

For a term the model may not know, or one the industry uses inconsistently, keep the term and translate it into the visible change:

Rack focus: shift focus smoothly from the leaves in the foreground to the
person in the background. The leaves gradually blur while the person's face
changes from soft to sharp.

Aperture, focal length, and shutter values are allowed, but the visible result you want is almost always the clearer instruction.

Pre-flight checklist

Run this before you spend the generation:

  • Subject and primary action are stated plainly.
  • Every reference says what to use and what not to use.
  • Every distinct character, product, and prop is named and bound to one reference.
  • References are picked per scene, not all forced into every shot.
  • Each stage of a long video has one main change and a clear end state.
  • Character count, clothing, prop ownership, and spatial relationships stay consistent.
  • Edits name the sole master, the scope, and the content to preserve.
  • Emotions and unusual camera terms come with observable cues.
  • First and last images share an aspect ratio, and each keyframe has one role.
  • Extensions have a checked boundary frame, motion trend, and audio continuity.

What Seedance 2.5 will not do precisely

Worth knowing before you plan a shot around it:

  • Timestamps allocate time to events. They are not frame-accurate edit points.
  • An edit raises the odds that events line up with the source clip; it cannot guarantee frame-by-frame overlap.
  • A seamless transition aims at visual and audio continuity, not pixel-identical preservation of both source clips.
  • Subtitles, formulas, signage, and product specifications that must be exactly right are still a post-production job. Generate the shot, composite the text.

Weak vs strong prompts

Name the camera, the motion over time, and the role of each reference rather than leaving them to chance.

FocusWeakStrong
CameraA city skyline at duskLow aerial gliding over a skyline at dusk, one unbroken push toward a single lit tower
Motion over timeA climber on a ridgeA climber hauls over the ridge, stands, and turns to the valley as the camera pulls back
ContinuityUse these references@Image 1 defines the detective's coat and face, do not use its background. @Image 2 defines the plaza's layout and light

Common mistakes

  • Describing a still: a video model needs motion over time, not a photograph in words.
  • Vague camera: "cinematic" tells the model nothing; name the shot and one move.
  • Cramming a sequence into one prompt without stages, so events collide instead of following each other.
  • Leaving references unlabeled, or labeling them as a group, so the model has to guess which one drives the scene.
  • Forgetting the exclusions, and inheriting a reference's background or bystanders along with the subject.

Simple pricing

Get started for free today, with the option to upgrade or cancel anytime.

Basic

$9/ month
billed as $0 per year

900 monthly credits

1 user only

All models

Workflows

Standard

$24/ month
billed as $0 per year

3200 monthly credits

1 user only

All models

Workflows

Pro

$45/ month
billed as $0 per year

6200 shared monthly credits

1 user

+ up to 4 more at extra cost

All models

Workflows

Pro Max

$170/ month
billed as $0 per year

24000 shared monthly credits

1 user

+ up to 9 more at extra cost

All models

Workflows

Enterprise

For higher limits

Custom

pricing and billing terms

High-volume credits
Custom seat limits
All models
Workflows
Pricing Gradient

Free

For playing around

$0

forever free

Up to 20 credits
1 user only
Limited models
Workflows

FAQs

How do I write a good Seedance 2.5 prompt?
Follow ByteDance's own formula: subject and action first, then scene and environment, visual style, camera movement or cuts, and audio. Write it as four short lines, one per element, and drop any line you do not need. Describe how the shot evolves over its full length rather than a single frozen frame, and give a specific camera move like a slow push-in instead of a vague "cinematic."
How do references work in Seedance 2.5?
You supply reference inputs, up to 50 per generation, such as an image of a character, product, or set, and the model carries that look through the clip and across shots. Describe in the prompt what each reference is for, so the model knows which one drives the subject, which sets the location, and which fixes the palette. Reusing the same references from cut to cut is what keeps a sequence consistent.
Can Seedance 2.5 make longer single-take videos?
Yes. Seedance 2.5 is built for longer native takes than most models, so a full reveal or a short scene can play out in one continuous shot instead of several clips stitched together. Write the prompt as one evolving motion, describing how the subject and camera move across the whole length of the take.
Can I edit a clip without regenerating it in Seedance 2.5?
Yes. For a targeted change to an existing clip, adjust that element in place rather than re-rolling the whole shot. Editing one part keeps the rest of the frame stable, which is faster than a full regeneration and avoids losing a take you already like.
What do the brackets mean in a Seedance 2.5 prompt?
Seedance 2.5 reads four bracket types as labels for the kind of sound or text you mean. Round brackets mark music, angle brackets mark a sound effect, curly braces mark spoken dialogue, and full-width square brackets mark an on-screen subtitle. They are optional, since plain language works too, but they stop a line of dialogue being drawn as a caption or a sound effect being read aloud.
How do I control timing in a Seedance 2.5 video?
Divide the video into consecutive stages and give each stage one main change plus a visible end state. Reach for timestamps only for a critical handoff, entrance, or transition, using a time range to budget pacing, an exact second for one key beat, or a relative delay between two events. Timestamps allocate time to events rather than acting as frame-accurate edit points.
What resolution can Seedance 2.5 output?
ByteDance's Seedance 2.5 guide doesn't publish an output resolution yet. It documents reference inputs of up to 4K for images and 480p to 4K for video clips, and the output specification will be confirmed at launch. See the Seedance 2.5 model page, which will update once the number is official.
How do I use Seedance 2.5?
Switch the prompt bar to video, write the shot as a short brief covering subject, action, scene, style, camera, and audio, and attach any reference images for a character, product, or set that has to stay consistent. Name what each reference is for and what to ignore in it. Run the prompt, then refine the keeper with in-clip editing rather than regenerating the whole shot.