Seedance 2.5 features and capabilities
Seedance 2.5 is ByteDance's next-generation video model. It extends the Seedance 2.0 line with longer native takes, many more references, and native audio across 10+ languages, while keeping the AI camera control the line is known for.
Seedance 2.5 is rolling out, so some details may change before or at its full launch, which ByteDance expects on August 7, 2026.
| Feature | What it does | Best for |
|---|---|---|
| Longer single-take video | Holds one continuous shot far past a few-second beat, up to 30 seconds | Ad spots, short scenes, long reveals |
| Reference-led continuity | Takes up to 50 multimodal references to lock a character, set, and palette | Series, multi-shot sequences, brand work |
| In-clip editing | Changes one element of an existing clip without a full re-render | Fixing a take, targeted revisions |
| AI camera control with audio | Directs the camera in plain language, with native audio in the same pass | Directed shots, sound-on scenes |
| Audio-only reference | A voice, music, or sound-effect track can drive pacing, beat-matching, and lip-sync | Music videos, lip-synced dialogue, localized voiceover |
Longer single-take video
The headline change is length. Seedance 2.5 holds a single continuous take far past the few-second clips most models produce, so a full reveal or a short scene can play out in one shot instead of several stitched together. You write the prompt as one evolving motion, describing how the subject and camera move across the whole take rather than a single frozen frame.
Reference-led continuity
Seedance 2.5 leans on multimodal references to keep a look steady. Supply an image of a character, product, or set, and the model carries that look through the clip and from shot to shot. It accepts up to 50 reference inputs per generation (30 images, 10 video clips, and 10 audio), spanning images, video clips, audio, and even simple 3D white models, so a team can lock a character, a location, and a palette together, which is what keeps a sequence recognizable rather than drifting scene to scene.
In-clip editing and extension
Editing is targeted rather than a full re-roll. For a single change to an existing clip, adjust that element in place and the rest of the frame stays put, which is faster than regenerating the whole shot and avoids losing a take you already like. You can point at a moment by timestamp, so a change lands only where you want it.
| Operation | What it does | Example |
|---|---|---|
| Instruction edit | Adds, removes, or changes content by text, optionally at a set timestamp | Replace the scene from 0:02 to 0:05, keep the motion and rhythm |
| Reference-image edit | Swaps an element to match a supplied image | Replace the jacket with the one in the reference, keep the rest |
| Add a subject | Drops a person, prop, or effect into the frame, rest unchanged | Add an energy bow into the character's hands |
| Remove a subject | Erases an object, watermark, or logo, then fills the gap | Remove the drone and its track, inpaint the sky |
| Audio edit | Replaces the music, adds effects, or changes the voice | Swap the background music, keep the visuals |
| Extend a clip | Continues forward, fills the moment before, or bridges two clips | Extend six seconds forward as the sky darkens |
Extension is high-fidelity: it can continue a clip past its last frame (forward), generate the moment before its first frame (backward), or build a transition that connects two clips, keeping character, rhythm, and style intact. You set how many seconds to add.
AI camera control with audio
Camera moves are directed in plain language, a slow push-in, a low tracking shot, a pull-back, so motion reads as an intentional camera rather than drift. Native audio generates in the same pass, so a scene can come back with room tone or effects already in place instead of a silent clip.
Native audio and 10+ languages
Sound generates in the same pass, so a scene can come back with voice, music, or effects already in place. New in 2.5, an audio-only reference lets a single voice, music, or sound-effect track drive the pacing, beat-matching, and lip-sync of the shot. Seedance 2.5 also generates natively in 10+ languages and follows instructions more closely, so you can describe a scene, and localize its voiceover and titles, in your own language.
Seedance 2.5 use cases
Long single-take establishing shots
One unbroken move across a landscape, held far past a few-second beat. A long native take carries a whole reveal without a cut, so an aerial reads as a single continuous camera rather than stitched clips.
Detailed close-up craft
Hands at work, jewelry, and mechanisms hold their fine detail through the move. Reference-led generation keeps the object consistent frame to frame, so a close inspection shot stays crisp instead of smearing.
Sports and athletic motion
Fast bodies and shifting weight stay coherent across the clip, with motion that tracks the action rather than blurring it. A dawn sprint reads with real timing and follow-through.
Cinematic sci-fi scenes
Big sets, atmosphere, and volumetric light give a shot scale. Camera control moves through the space with intent, so a hangar interior feels staged for a scene rather than a static render.
Travel and documentary
Wide vistas and slow reveals finish clean at high resolution, ready for a large screen. A desert crossing holds its detail from foreground grain to the far horizon.
Surreal, imaginative concepts
Impossible scenes hold together because the subject stays consistent through the shot. A whale drifting past a diver keeps its scale and weight, so the concept lands instead of falling apart mid-move.
How to get the best out of Seedance 2.5
Seedance 2.5 rewards a clear shot brief and a habit of leaning on references for continuity. A few practices carry most of the quality:
- Write for one continuous take. Describe how the subject and camera evolve across the whole length of the shot, not a single instant.
- Name the camera move. A specific "slow push-in" or "low tracking shot" reads as intent, where "cinematic" tells the model nothing.
- Lock the subject with a reference. Supply a clear image of a character, product, or set so the look carries through the clip instead of drifting.
- Say what each reference is for. Name the character, the set, and the palette so the model knows which reference drives which part of the scene.
- Reuse references across shots. Carrying one set of references from cut to cut is what keeps a sequence consistent.
- Stage complex blocking with a 3D whitebox. A simple white-model scene passed as a reference lets you lay out the set and the camera path before you spend a render on it.
- Edit in place, don't re-roll. For a targeted fix to an existing clip, change that element rather than regenerating the whole shot.
- Drive rhythm with an audio-only reference. Attach a music or voice track and the model can match the pacing, beats, and lip-sync to it.
Choosing references for stable results
More references is not always better; a focused set is what holds quality:
- With image references, keep the main subjects to 8 or fewer, and push to 9–12 only when you have to.
- With a video or audio reference, keep subjects to 5 or fewer, and aim for a 5 to 10 second clip: long enough to read, short enough to stay clean. Video references total 30 seconds across all of them, and so does audio.
- When editing an existing video, keep the source clip under 20 seconds and pair it with 1 to 5 reference images, pushing to 6–8 only when the edit needs them.
- With five or fewer subjects, single or multiple views both work; past five, prefer single-view images, or upload separate angles rather than one combined image. Separate angle images hold up better than several views collaged into one picture.
- Prioritize in this order: core characters, then key products or props, then the scene, then overall style.
For the full capability list and specifications, see the Seedance 2.5 model page.
Seedance 2.5 prompt guide
A strong video prompt reads like a short shot brief, not a caption, so the model has a subject, a motion, and a camera to work with rather than a still frame in words. ByteDance's own prompt guide sets out one formula for this, and everything else on this page builds on it.
The Seedance 2.5 prompt formula
Six elements, in this order, and you drop any you do not need:
| Element | What it covers | Required |
|---|---|---|
| Subject and action | Who or what is in frame, and what they do | Yes, this is the foundation |
| Scene and environment | Location, time, weather, spatial relationships, background | Optional |
| Visual style | Lighting, color, materials, texture, overall mood | Optional |
| Camera movement or cuts | Shot size, angle, movement, focus subject, transitions | Optional |
| Audio | Dialogue, voice, ambience, sound effects, music | Optional |
Written out, that is four short lines:
[Subject] performs [primary action or event] in [scene and environment].
The visuals feature [visual style].
Use [shot size, camera angle, camera movement, or cuts].
Audio includes [dialogue, ambience, sound effects, or music].
A worked example:
A ceramic artist finishes a pale blue cup in a studio at dawn, lifts it from
the wheel, and places it in the center of a wooden shelf.
Soft morning light enters through the window. The wet clay has a delicate
sheen, and the workbench remains tidy.
Begin with a medium shot of the wheel-throwing process, slowly push in toward
the cup's surface texture, then cut to a frontal view of the shelf.
Retain the low hum of the pottery wheel, the friction of clay, and subtle
indoor ambience.
Aspect ratio, duration, and the other generation settings stay out of the prompt. You set those on the generation page, not in the text.
Audio, dialogue, and subtitle syntax
Plain language works, but four bracket types tell Seedance 2.5 exactly which kind of sound or text you mean. They are the cheapest way to stop a line of dialogue being drawn on screen as a caption, or a sound effect being spoken aloud.
| Content | Bracket | Example |
|---|---|---|
| Music | ( ) | (Soft, rhythmic piano music plays in the background) |
| Sound effect | < > | <A bell rings in the distance> |
| Dialogue | { } | {Hello, welcome back.} |
| Subtitle | 【 】 | 【Chapter One: Departure】 |
Name the language before any line that is not in Chinese, and name it again if the model keeps reaching for the wrong one. The pattern is language, then regional variety or accent, then delivery style, then speaker, then the line:
Dialogue language: American English. The girl says in natural, conversational
American English: {I thought you weren't coming.}
That last part matters more than it looks. Naming a regional variety, "authentic Los Angeles English" rather than just "English", is what gets you a specific delivery instead of a neutral read.
Naming references and defining their roles
References are where Seedance 2.5 pulls away from earlier models, and where most weak prompts fall down. Uploading an image is only half the instruction. The prompt has to say what that image contributes, and what to ignore in it:
@Image 1 defines the ceramic artist's facial features, hairstyle, and dark
green apron. Do not use the image background.
@Image 2 defines the wooden workbench, window placement, and morning light of
the pottery studio. Do not use the people in the image.
@Video 1 defines the pacing of throwing clay with both hands, lifting the cup,
and placing it down. Do not use the person's identity, clothing, or scene.
Three rules carry most of the reliability here:
- Bind one reference to one subject at a time. Write out each mapping. A line like "@Images 1 through 4 define four characters respectively" never says which image is which character.
- Add the exclusions. Backgrounds, bystanders, and compositions ride along from a reference unless you rule them out.
- Say when several images are one thing. Four angles of the same lamp need "all four images define one folding desk lamp, and the output must contain only one lamp throughout", or you get four lamps.
When a reference video already carries the motion and camera work, name only the attributes to inherit. Restating the action in words fights the reference you attached.
Directing many references at once
Past a handful of references, the goal stops being description and becomes casting. Work in this order: define each material's role, map subjects, group by type, build profiles for the important ones, then pick references per scene.
Grouping keeps a long list readable:
[Characters]
<Conservator> corresponds to @Image 1. Use only the appearance,
hairstyle, and clothing.
<Registrar> corresponds to @Image 2. Use only the appearance,
hairstyle, and clothing.
Do not interchange their appearances, clothing, actions, or dialogue.
[Props]
<Sample Case> corresponds to @Image 5 and belongs only to <Conservator>.
[Scenes]
<Conservation Lab> references @Image 7. Use only the space,
materials, and lighting.
For a character who recurs across scenes, write one profile that gathers every reference attached to them, including a "do not use" line. Then select per scene, so each shot draws only the references it needs:
Scene 1 | Inspection in the Conservation Lab
Use: <Conservator>, <Sample Case>, <Conservation Lab>, and the
case-opening motion from @Video 1.
Event: <Conservator> opens <Sample Case> at the workbench and
inspects the sample inside.
End state: <Conservator> remains on the inner side of the workbench,
<Sample Case> beside their right hand.
Fifty references is a casting pool, not a shopping list. The point is helping the model pick the right materials for the shot in front of it, not getting everything on screen at once.
Longer, edited, and extended Seedance 2.5 shots
Length is what Seedance 2.5 is built for, and a 30-second take needs a different prompt shape than a five-second one. So do the edit and extension modes, which run on the clip you already have rather than a blank frame.
Writing a 30-second Seedance 2.5 video
Split the story into consecutive stages. Give each stage exactly one main change, and state what should be visible on screen when it ends. The end state is what carries continuity from one stage to the next:
[Generation Goal]
Generate an instructional video showing a flower shop's order-packing process.
[Stage 1]
Initial state: <Florist> stands behind the workbench. Loose stems,
scissors, and wrapping paper lie on the tabletop.
Primary event: <Florist> arranges the stems and trims them to length.
End state: <Florist> holds the bouquet in the left hand, scissors
back on the right of the workbench.
[Stage 2]
Continue from the previous stage: both characters keep the same
identities and clothing.
Primary event: <Store Assistant> unfolds the wrapping paper,
<Florist> places the bouquet inside and ties it.
End state: the wrapped bouquet lies flat in the center of the
workbench, bow facing camera.
[Maintain Consistency]
Keep character identities, clothing, prop ownership, and workbench
orientation consistent.
Timing and pacing
Stages are the default. Reach for timestamps only when a handoff, entrance, transition, or beat has to land at a particular moment.
| Pattern | Use it for | Example |
|---|---|---|
| Time range | Budgeting pacing across the clip | 0–3 seconds… 3–7 seconds… 7–12 seconds… |
| Exact time point | One critical beat | At 5 seconds, the camera whip-pans left and completes the transition |
| Relative timing | A delay between two events | Three seconds after the character presses the button, the lights go out |
Keep ranges consecutive and non-overlapping. They are time budgets rather than edit points, so an action can land slightly either side of a boundary. Too little content in a range gives the model room to wander; too much causes rushed cuts or dropped events. Asking for three actions in one second does not work.
Editing an existing clip
An edit is a patch, not a new prompt. Name the source clip as the sole master, state the scope of the change, and list what must not move:
[Edit Goal]
Edit @Video 1. Only from 4-7 seconds, change the cool blue light on
the right wall to warm orange light.
[Source Video Role]
@Video 1 is the sole editing master. It defines the character, room
layout, actions, composition, camera movement, audio, and event order.
[Edit Scope]
Change only the light color on the right wall and the area it
illuminates. Allow skin tone to respond naturally.
[Content to Preserve]
Keep the character's identity, clothing, expression, position, motion,
room structure, camera movement, dialogue, and ambience from @Video 1.
Swapping a subject adds one more block, and it is the block people forget. The replacement has to inherit the original's timeline: every appearance, movement, occlusion, and exit, with the same timing, path, and speed. Without it you get the new object in the right frame and the wrong rhythm. Background replacement works the same way, scoped to everything outside the subject's silhouette.
Audio edits are separate from visual ones. Name the speaker or sound category, the change, and what stays: removing the music while keeping dialogue, lip sync, ambience, and effects is a single instruction.
Extending a clip forward or backward
Extension bolts new footage onto a boundary frame, so describe that frame before you describe the new action. Forward, the extension's first frame continues from the source's last frame. Backward, the extension's last frame has to arrive at the source's first frame:
@Video 1 is the source video to extend forward.
Extend @Video 1 forward. The first frame of the extended segment directly
continues from the last frame of @Video 1. Maintain the same locked-off medium
shot, the paper airplane's position and orientation, the classroom-window
background, the afternoon lighting, and its movement toward frame right.
Then, the paper airplane continues gliding right and exits the frame while the
curtain beside the window sways slightly.
Backward extensions need one extra guard. Say which materials belong only to the original clip, or characters and props that should arrive later turn up early in the segment you just generated. And treat "connect to the source video" as insufficient on its own: spell out the source's first frame as the explicit end state, or the image keeps changing after it has already arrived.
Locked settings you cannot override
Editing, extension, and first-frame generation each lock some parameters to the material you fed them.
| Task | Aspect ratio | Duration |
|---|---|---|
| Video editing | Inherited from the input video, cannot be set | Inherited, cannot be set, within about 0.3 seconds |
| First or first-and-last frame | Inherited from the first image | Can be set |
| Video extension | Inherited from the input video, cannot be set | Can be set |
The practical catch is on first-and-last-frame work: give both images the same aspect ratio, or the last frame gets stretched to match the first.
Advanced Seedance 2.5 direction
Keyframes, storyboards, and 3D blockouts
You can hand Seedance 2.5 the structure of a shot rather than describing it. Four ways, in rising order of control:
- First and last frames. Say in the first line that one image is the first frame and another is the last, then describe the single continuous action between them. Describe each anchor separately; a combined "these two are the first and last frames" does not bind either one.
- Multi-keyframe sequences. Open with "use @Image 1 through @Image N as keyframes in this order", then give the key state each image represents. Separate images align better than several frames collaged into one grid, and they control stage order rather than reproducing every frame.
- Storyboard grids. A grid communicates shot order and rough composition, not exact detail. Keep it under about 15 panels, use clean line art, minimize text labels, state the reading order, and rule out the grid's own style so the line art does not end up in the render.
- 3D blockouts. A coarse blockout carries timing: paths, blocking, camera movement, cut points, lighting changes, and sound rhythm. Map each grey shape to its final subject by name. A fine blockout already has complete structures, so it is for re-rendering materials, characters, scenes, and style while the structure and camera work hold. Strip path lines, coordinate axes, and camera frustums out of a fine blockout first, or they get rendered too.
Directing performance and emotion
Emotion words set a direction but leave the acting open. What tightens it is naming what a viewer would actually see or hear: eye movement, brow tension, mouth movement, breathing, gaze direction, hand movement. Two to four cues is usually enough for a single emotional turn, and you never need to list every facial detail.
The default shape is a trigger, an immediate reaction, a gradual change, then the settled expression:
The overall emotion shifts from [starting emotion] to [ending emotion].
After [triggering event], [subject] first shows [immediate observable reaction].
Then, [eyes, brows, mouth, breathing, gaze, or hand movement] gradually [changes].
Finally, [subject] expresses [target emotion] through [outward behavior].
For an emotion that turns several times, chain it to events instead: each new thing the character hears or sees moves the performance one step.
Camera language Seedance 2.5 understands
Standard vocabulary can go straight into the prompt: shot sizes from extreme wide to extreme close-up; push in, pull out, pan, lateral move, follow, orbit, dive, dolly out, tilt up, handheld shake; low angle, overhead, first person.
Popular techniques work too, as long as you say which subject they apply to and where the move starts and ends. A one-take shot needs the spaces it passes through in order; a dolly zoom needs the subject size to hold and which way the background should travel; bullet time needs the action to freeze and the orbit direction; a speed ramp needs where the action accelerates and where it settles.
For a term the model may not know, or one the industry uses inconsistently, keep the term and translate it into the visible change:
Rack focus: shift focus smoothly from the leaves in the foreground to the
person in the background. The leaves gradually blur while the person's face
changes from soft to sharp.
Aperture, focal length, and shutter values are allowed, but the visible result you want is almost always the clearer instruction.
Pre-flight checklist
Run this before you spend the generation:
- Subject and primary action are stated plainly.
- Every reference says what to use and what not to use.
- Every distinct character, product, and prop is named and bound to one reference.
- References are picked per scene, not all forced into every shot.
- Each stage of a long video has one main change and a clear end state.
- Character count, clothing, prop ownership, and spatial relationships stay consistent.
- Edits name the sole master, the scope, and the content to preserve.
- Emotions and unusual camera terms come with observable cues.
- First and last images share an aspect ratio, and each keyframe has one role.
- Extensions have a checked boundary frame, motion trend, and audio continuity.
What Seedance 2.5 will not do precisely
Worth knowing before you plan a shot around it:
- Timestamps allocate time to events. They are not frame-accurate edit points.
- An edit raises the odds that events line up with the source clip; it cannot guarantee frame-by-frame overlap.
- A seamless transition aims at visual and audio continuity, not pixel-identical preservation of both source clips.
- Subtitles, formulas, signage, and product specifications that must be exactly right are still a post-production job. Generate the shot, composite the text.
Weak vs strong prompts
Name the camera, the motion over time, and the role of each reference rather than leaving them to chance.
| Focus | Weak | Strong |
|---|---|---|
| Camera | A city skyline at dusk | Low aerial gliding over a skyline at dusk, one unbroken push toward a single lit tower |
| Motion over time | A climber on a ridge | A climber hauls over the ridge, stands, and turns to the valley as the camera pulls back |
| Continuity | Use these references | @Image 1 defines the detective's coat and face, do not use its background. @Image 2 defines the plaza's layout and light |
Common mistakes
- Describing a still: a video model needs motion over time, not a photograph in words.
- Vague camera: "cinematic" tells the model nothing; name the shot and one move.
- Cramming a sequence into one prompt without stages, so events collide instead of following each other.
- Leaving references unlabeled, or labeling them as a group, so the model has to guess which one drives the scene.
- Forgetting the exclusions, and inheriting a reference's background or bystanders along with the subject.

