What is Seedance 2.0?
Seedance 2.0 is ByteDance's advanced multimodal AI video model, combining images, videos, audio, and text inputs for unprecedented creative control. This complete guide compares Seedance to Kling, Veo, and Sora, and shows professionals how to master multimodal video workflows on Morphic.
Unlike traditional text-to-video models that rely solely on written prompts, Seedance 2.0 enables you to show the AI exactly what you want through visual and audio references. Upload reference images to define style and composition, use video clips to demonstrate desired camera movements or actions, add audio to establish mood and rhythm, and combine everything with detailed text prompts for precise creative direction.
Two companion guides go deeper on specific problems: cinematic Seedance 2.0 videos for the craft side, and what to do when Seedance 2.0 prompts get flagged for generations that come back rejected.
Why Seedance 2.0 for professional video creation
Seedance 2.0 addresses the fundamental limitation of AI video generation: the gap between description and vision. Instead of trying to describe complex camera movements, character details, or visual effects in words, you can provide direct examples. This multimodal approach delivers:
- Precise visual control through image references
- Accurate motion replication via video references
- Rhythm and mood synchronization with audio integration
- Consistent character and style across multiple shots
- Complex scene transitions that maintain continuity
The model excels at understanding and combining multiple reference types simultaneously, making it particularly valuable for commercial production, content creation, and professional video workflows.
Seedance 2.0 vs Kling vs Veo vs Sora: Feature comparison
When evaluating AI video generation tools, understanding the specific capabilities of each platform helps inform the right choice for your workflow. Here's how Seedance 2.0 compares to leading alternatives:
| Feature | Seedance 2.0 | Kling 3.0 | Veo 3.1 | Sora |
|---|---|---|---|---|
| Multimodal input support | Images, videos, audio, text | Images, videos, audio, text | Images, text | Images, text |
| Maximum video duration | Up to 15 seconds | Up to 15 seconds | Up to 8 seconds (extendable to 60+ seconds) | Up to 60 seconds |
| Audio integration | Direct audio upload and reference | Native audio with lip-sync, multi-language dialogue | Native audio with sound effects and dialogue | Text-to-audio only |
| Video reference capability | Full motion and camera replication | Full motion and camera replication with AI director | Style transfer and reference images (up to 3) | Limited |
| Maximum output resolution | 4K (10-bit) | 4K | 4K | 1080p |
Key differentiators:
Multimodal flexibility: Seedance 2.0 and Kling 3.0 both offer comprehensive multimodal support including direct video and audio file uploads. Veo 3.1 supports image references (up to 3) but audio is generated rather than referenced. Sora remains primarily text and image-based.
Video reference depth: Seedance 2.0 and Kling 3.0 excel at replicating complex camera movements, choreography, and special effects from reference footage. Kling 3.0's "AI Director" feature automates multi-shot scene composition. Veo 3.1 focuses on image-to-video with strong character consistency but less emphasis on video-to-video motion replication.
Audio capabilities: Seedance 2.0 allows direct audio file upload for precise mood control and beat synchronization. Kling 3.0 generates native multi-language audio with accurate lip-sync across 5 languages. Veo 3.1 generates audio natively but doesn't accept audio file references. Sora generates audio from text descriptions only.
Duration and extension: While Sora offers the longest single generations (up to 60 seconds), Veo 3.1's extension feature allows chaining clips beyond 60 seconds. Seedance 2.0 and Kling 3.0 both support 15-second generations with extension capabilities.
Resolution and quality: Seedance 2.0 outputs up to 4K with 10-bit encoding, which preserves colour gradation for HDR and grading work. That puts it level with Kling 3.0 and Veo 3.1 on raw resolution, so the deciding factor is usually reference control rather than pixel count.
Editing existing footage: Seedance 2.0 treats an uploaded video as something you can edit, not just imitate. You can add, remove, or swap elements inside a clip, extend it forward or backward in time, and stitch separate clips into one continuous piece. Editing is a first-class task here rather than a bolt-on, which is what makes it practical to keep footage a client has already approved.
Which Seedance 2.0 model to use
Seedance 2.0 is a family of three models that share the same feature set and differ on quality and cost:
| Model | Best for | Maximum resolution |
|---|---|---|
| Seedance 2.0 | Highest generation quality; the only one that outputs 4K | 4K |
| Seedance 2.0 Fast | Balancing speed and cost when top-tier quality is not required | 720p |
| Seedance 2.0 Mini | Best cost efficiency, high-volume iteration | 720p |
A practical pattern: draft on Mini or Fast until the prompt and references are behaving, then run the final pass on full Seedance 2.0.
Information accurate as of July 2026. Features and availability subject to change.
Key features and capabilities of Seedance 2.0
Multimodal input system
Seedance 2.0 accepts four distinct input types that work in combination:
Image Inputs (Up to 9 images)
- Define visual style and aesthetic direction
- Establish character appearances and maintain consistency
- Set scene composition and framing
- Specify product details for accurate reproduction
- Control lighting, color grading, and atmosphere
Video Inputs (Up to 3 clips, maximum 15 seconds combined)
- Reference specific camera movements and cinematography
- Replicate motion patterns and choreography
- Copy scene transitions and editing rhythms
- Demonstrate special effects and visual techniques
- Show character actions and interactions
Audio Inputs (MP3 or WAV, up to 3 files, each 2-15 seconds, maximum 15 seconds combined)
- Set mood and emotional tone through music
- Control pacing with rhythm and beat structure
- Add specific sound effects or ambient audio
- Match voice characteristics for dialogue
- Synchronize visual changes to audio cues
Text Prompts (Natural language)
- Guide narrative and story progression
- Specify actions and movements not shown in references
- Describe scene transitions and timing
- Clarify how references should be applied
- Add details beyond what visual references show
On limits: there is no combined file cap. The ceilings are per type (up to 9 images, 3 video clips, 3 audio files), and the real constraint is not the ceiling but your restraint: ByteDance recommends using far fewer references than you are allowed. One combination is genuinely unsupported, though. Text plus audio alone, and audio alone, will not generate. Every job needs text, and audio only ever works alongside a visual reference.
Reference capability architecture
The core innovation in Seedance 2.0 is its reference understanding system. Rather than treating inputs as simple style guides, the model analyzes and extracts specific elements from each reference:
From Images: Composition structure, character features, object details, lighting setup, color relationships, spatial arrangement, style characteristics
From Videos: Camera motion paths, movement speed and acceleration, shot framing changes, subject actions and timing, special effect implementation, transition techniques
From Audio: Rhythm and beat patterns, tonal mood and atmosphere, volume dynamics, sound effect timing, voice characteristics
This granular understanding allows you to specify exactly which aspects of each reference should influence the generation, creating precise control over the final output.
Core generation quality improvements
Beyond multimodal capabilities, Seedance 2.0 delivers foundational enhancements:
Realistic Physical Dynamics: Objects and characters move with authentic physics. Clothing drapes naturally, liquids flow convincingly, and interactions between elements follow real-world rules.
Smooth Motion Performance: Continuous action flows without jarring transitions or morphing artifacts. Complex multi-step movements maintain consistency throughout execution.
Precise Prompt Understanding: The model accurately interprets detailed instructions, including spatial relationships ("in the background behind"), ordered shot sequences, and complex multi-subject scenarios. Precise timecodes are the exception, and are covered below.
Consistent Style Retention: Visual characteristics established at the start of a generation remain stable throughout. Character appearances, lighting conditions, and artistic style don't drift as the scene progresses.
Complex Action Execution: Handles challenging sequences like fight choreography, detailed hand movements, facial expressions during speech, and coordinated multi-character interactions.
Ready to experience multimodal control? Start creating with Seedance 2.0 on Morphic →
Technical specifications
| Parameter | Specification |
|---|---|
| Generation duration | 4-15 seconds |
| Output resolution | 480p, 720p, 1080p, 4K (Seedance 2.0). Fast and Mini top out at 720p |
| Aspect ratios | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, plus an adaptive setting that matches your reference image |
| Output format | MP4. 4K is encoded H.265 (HEVC), 10-bit |
| Audio output | Integrated sound effects, dialogue, and background music generation |
| Image inputs | JPEG, PNG, WEBP, BMP, TIFF, GIF, HEIC, HEIF. Up to 30MB each |
| Video inputs | MP4, MOV (H.264 or H.265), 24-60fps, 2-15s per clip, up to 200MB each |
| Audio inputs | MP3, WAV. 2-15s per file, up to 15MB each |
Two notes on the output settings. There is no 2.35:1 option, so for a scope look, choose 21:9 rather than asking for 2.35:1 in the prompt. And because 4K uses H.265, some browsers and players will not preview it directly even though the file is fine.
Understanding Seedance 2.0 input specifications
File count and duration limitations
To optimize generation quality while managing computational resources, Seedance 2.0 implements specific input constraints:
Individual File Type Limits:
- Images: Maximum 9 files
- Videos: Maximum 3 clips
- Audio: Maximum 3 files
Combined Duration Limits:
- Video references: 15 seconds total across all clips
- Audio references: 15 seconds total across all files
Per-clip minimums (easy to miss):
- Every reference video clip must be at least 2 seconds long
- Every reference audio file must be at least 2 seconds long
- Reference video must run at 24 to 60 fps
There is no combined file cap. Nine images plus three videos plus three audio files is a legal request. It is also a bad one, for the reasons in the next section.
Unsupported combination: text plus audio, and audio on its own, will not generate. Audio always needs a visual reference alongside it.
Generated output duration: 4-15 seconds.
Do not fill the slots you are given
The limits above are ceilings, not targets. ByteDance's own guidance is explicit that using the full asset allowance degrades results: with too many references, the model struggles to judge which features take priority, and you get style conflicts, blurry subject identification, and output that drifts from what you asked for.
The recommended configuration is 4 to 5 assets in total. Each one should do a distinct job:
- Character anchoring (1-2 images): lock the character's appearance
- Scene tone-setting (1 image): lock the environment and style
- Camera movement (1 video): lock the shot language and rhythm
- Rhythm and atmosphere (1 audio): control emotion and timbre
If two references are competing for the same job, you have one too many. Uploading four images of the same character does not make the model four times more confident about that character. It makes the model less certain about which image to believe.
Practical example: for a 15-second commercial requiring specific product appearance, dynamic camera work, and upbeat music, five assets is a complete brief:
- 2 images: product from two angles
- 1 image: desired colour grading and lighting style
- 1 video: camera movement reference
- 1 audio: music track for pacing
Everything else, including the environment, the mood, and the pacing, belongs in the text prompt.
A hard ceiling on characters: when more than four reference people are involved, output stability drops sharply. You start seeing the wrong number of people in frame, or duplicates. If you need a crowd of six, generate them as two grouped images of three first, then use those images as the reference for the video.
Input quality guidelines
For Image References:
- Use clear, well-lit photographs when accuracy matters
- Higher resolution provides better detail reproduction
- Multiple angles help for products and objects. They actively hurt for faces: see the headshot rule below, because a multi-angle character sheet is read as several different people
- Avoid heavily compressed or low-quality images
For Video References:
- Ensure the specific element you want to reference is clearly visible
- Shorter clips focused on one aspect work better than longer clips with multiple elements
- Higher quality video improves motion understanding
- Trim videos to show only the relevant section
For Audio References:
- Use clean audio files without background noise when possible
- Ensure audio clearly demonstrates the rhythm or mood you want
- Match approximate duration to your target video length
- Consider using audio from video files if it serves multiple purposes
How to use Seedance 2.0 multimodal references
Seedance 2.0 is accessible through Morphic, which provides an interface for uploading references and writing prompts. The system uses an @ mention structure to specify how each uploaded file should be used in generation.
First, know which of the three tasks you are asking for
Before writing anything, decide which job you are giving the model, because the phrasing differs and getting it wrong is the most common way to waste a generation.
- Multimodal reference: pull elements out of your assets (a subject, a style, a camera move, a voice) to make a brand-new video.
- Video editing: change something inside an existing video. Anything you do not mention stays as it was.
- Video extension: continue an existing video forward or backward in time, keeping the style, subject, and narrative intact.
The trap: for editing and extension tasks, refer to the clip as @Video 1, never as "reference @Video 1". The word reference pushes the model to classify the job as a reference task, so instead of editing your clip it generates a new video that merely resembles it. Say "Strictly edit @Video 1" or "Extend @Video 1 backward", not "Reference @Video 1 and edit it".
Define your subject before you reference it
An uploaded image usually contains more than one thing. Naming the subject explicitly is the step most people skip, and it is the one that prevents the model from latching onto the wrong element.
Pick 2 to 3 stable, static features (clothing, hairstyle, category) that uniquely identify the subject, then bind them to the asset:
Define [core features] in @Image 1 as [Subject name]
For example: Define the woman in the red dress and straw hat in @Image 1 as Subject 1. From that point on, refer to her only as Subject 1. Use the same label every single time she appears. If you switch between "the woman", "she", and "Subject 1", the model can lose track of who you mean.
For quick jobs you can skip the definition and bind inline instead, repeating the pairing each time you mention her: Subject 1@Image 1.
Two rules that matter here. Keep the description short and free of contradictions, because two conflicting features on the same subject will confuse the binding. And where a spatial relationship is easier to show than to say, show it in a reference image rather than describing it in text.
The @ reference system
After uploading your materials to Morphic, you reference them in your prompt using the @ symbol followed by the file identifier (Image 1, Video 1, Audio 1, etc.). The key is explicitly stating what purpose each reference serves.
Basic Reference Structure:
@[Material Type + Number] as/for [specific purpose], [additional context]
Clear vs Unclear Referencing:
Unclear: "Use @Image 1 and @Video 1 to make a video"
Clear: "@Image 1 as the opening frame showing the character's face, reference the camera push-in movement from @Video 1, use @Audio 1 for background music to establish an upbeat mood"
Put your most important asset first. The earlier an asset appears in the prompt, the more weight it carries. If the character's face has to be exactly right, lead with it.
Punctuation that tells the model what kind of information it is reading
Seedance 2.0 reads four bracket types as type signals. Using them is a cheap way to stop dialogue being rendered as an on-screen caption, or a sound effect being spoken aloud.
| Information type | Bracket | Example |
|---|---|---|
| Music | ( ) | (fast-paced rock music plays in the background) |
| Sound effect | < > | <a dog barks in the distance> |
| Dialogue | { } | {We should leave. Now.} |
| On-screen subtitles | 【 】 | 【Chapter One: Departure】 |
If a line of dialogue is in a language other than English, name the language before the braces: says in Japanese {こんにちは}. Keep one language per scene. Mixing languages in the same prompt, proper nouns aside, degrades pronunciation.
Writing effective multimodal prompts: The CRAFTS framework
A good Seedance 2.0 prompt is closer to an engineering instruction than a piece of copywriting. The model reads your text, images, video, and audio at the same time and splits them internally into a spatial layer (what is in the frame) and a temporal layer (how it changes over time). Your job is to feed both layers cleanly.
CRAFTS covers the six things a complete prompt specifies: Context, Reference, Action, Framing, Timeline, Style and constraints.
C - Context: Establish Scene and Environment Set the stage with location, time period, atmosphere, and overall setting. Include references to scene images here.
Example: "In a dimly lit jazz club at night, referencing the interior atmosphere from @Image 1"
R - Reference: Define Your Subjects and Bind Them to Assets Name each subject by 2 to 3 stable features, then state exactly what each asset contributes and what it does not.
Example: "Define the man in the grey suit in @Image 2 as the Pianist. @Image 2 for his face and build only, not his clothing. @Video 1 for the walking pace. @Audio 1 for the background jazz."
A - Action: Describe Character and Object Movements Detail what happens in the scene: character actions, object interactions, and event sequence. Two things separate action descriptions that work from ones that do not.
Go small and slow. Seedance 2.0 handles gentle, continuous, small-scale movement far more reliably than high-energy motion. Sprints, big jumps, and violent rolls are where limbs deform and physics breaks. Prefer "walks slowly", "gently raises a hand", "sits down with the motion". Specify the body part, plus the range, speed, and force: not "he moves his head" but "he turns his head quickly". Where one action leads into another, say how they connect, so the movement carries its own inertia: "uses the momentum of the turn to naturally raise a hand".
Externalise emotion into physical detail. The model cannot render "very sad". It can render what sadness does to a body. Replace the abstract label with the tell.
| Instead of | Describe this |
|---|---|
| Sadness | Head lowering, shoulders trembling slightly, eyes reddening, fingers unconsciously clutching the hem of a coat, tears welling but not falling |
| Joy | Corners of the mouth rising uncontrollably, brows and eyes relaxing, steps becoming light, humming without noticing, unable to resist spinning in place |
| Nervousness | Checking a watch repeatedly, fingers tapping the tabletop, breathing fast, eyes darting away, biting a fingernail |
| Anger | Both fists clenched, jaw tight, chest heaving, a hard stare, words squeezed out through gritted teeth |
| Relief | Letting out a long breath, tense shoulders dropping all at once, a faint smile returning to the face, looking up toward the distance |
Example: "The Pianist walks slowly across the room, stops at the bar, and picks up a glass. He takes a sip, and as he looks toward the door his jaw tightens and his fingers close hard around the glass."
F - Framing: Define Camera Work and Cinematography Specify shot types, camera movements, angles, and transitions using cinematic terminology. The model has a strong grasp of standard camera vocabulary, so use it plainly.
Example: "Start with a wide establishing shot, dolly in to a medium close-up as the character reaches the bar, then cut to an over-the-shoulder shot looking toward the door"
One camera move per shot. Asking a single shot to push in, orbit, and pan at once is the fastest way to destabilise the image. If you want three moves, you want three shots.
T - Timeline: Sequence your shots, not your seconds
This is the element most guides get wrong, including earlier versions of this one. The instinct is to timestamp everything: 0-4 seconds this, 4-8 seconds that. Seedance 2.0's support for precise timing is unstable, and forcing exact durations onto segments can actively break the generation.
Sequence the shots instead, and let the model find the pacing. Label them Shot 1, Shot 2, Shot 3, in the order events happen, and describe each one as: camera move, then subject action, then position, then sound.
Weak: "A man runs nervously down the street, and the scene feels very cinematic."
Strong:
- Shot 1: Side view of a street alley. The man starts to run, breathing fast.
- Shot 2: He knocks over a fruit stand. The camera shakes and pushes to a close-up of his frightened face.
- Shot 3: He climbs a low wall and disappears. The camera pulls back slowly and settles on the empty street.
Timestamps are not forbidden and you will see them in ByteDance's own sample prompts. Treat them as a hint the model may honour, not an instruction it must obey. If the pacing genuinely matters, the reliable lever is to generate fewer shots per clip, not to write tighter timecodes.
S - Style and constraints: Close the door on failure modes
The last block of your prompt is the one people leave out, and it is doing more work than they realise. It has three jobs.
Image quality: name the finish you want. "High definition, rich detail, filmic texture, natural colour, soft light."
Style: name the look, and name it even when you think the reference image already implies it. If your reference photo is realistic but you want animation, and you never say so, the output will drift back toward live action. Say "2D Japanese animation style" or "3D stylised CG" explicitly.
Constraints: these are the lines that stop the model producing artefacts nobody asked for.
- Keep it subtitle-free. Avoid generating any text or subtitles.
- Do not generate a logo.
- Do not generate a watermark.
Worth being straight about the limits here: constraint lines reduce the probability of stray subtitles and watermarks, they do not eliminate it. Two things that genuinely help. If a reference image or clip has text in it that you do not need, strip the text out of the asset before you upload it. And if your project allows it, generate in landscape, because spurious subtitles appear significantly less often in landscape than in portrait. You can always crop to vertical afterwards.
CRAFTS Example Prompt:
Context: A 1940s noir detective office at night, venetian blind shadows falling across the desk, with the lighting and atmosphere of @Image 1. Reference: @Image 2 for the detective's appearance (fedora, trench coat), @Video 1 for his slow, deliberate walking pace. Define the man in the fedora and trench coat in @Image 2 as the Detective. Shot 1: Wide shot of the full office. The Detective enters from frame left and walks toward his desk. <Footsteps on a wooden floor.> Shot 2: The camera tracks alongside him, then pushes in to a close-up as he lifts a photograph from the desk and studies it. Shot 3: Insert shot of the photograph in his hands. Shot 4: The camera pulls back to a medium shot. He sets the photo down and exhales heavily. Throughout: (a moody saxophone line from @Audio 1). Keep it subtitle-free. High-definition, filmic grain, warm practical light against deep shadow.
Image reference techniques
Setting visual style and aesthetic direction
Images establish the overall look and feel of your generation. Use them to define color palettes, lighting approaches, compositional style, and artistic treatment.
Create a rain-slicked night street scene with the visual style from @Image 1. Match its wet pavement reflections, shop-window glow, and moody blue-magenta colour grading. Use the vertical architecture composition from @Image 2. High definition, filmic grain, natural colour.
Maintaining character consistency across shots
When generating multiple videos featuring the same character, reference the same character image in each prompt to maintain appearance consistency.
Feature the woman from @Image 1 throughout this sequence, maintaining her exact facial features, hairstyle, and clothing. She starts in the outdoor setting from @Image 2, then the scene transitions to the indoor environment shown in @Image 3. Her appearance remains consistent across both locations.
Product showcase with accurate details
For commercial or product-focused content, use multiple angles and detail shots as references to ensure accurate reproduction.
Create a product showcase for the handbag in @Image 1. The side profile should match @Image 2, the surface texture and material details should reference @Image 3, and the hardware and clasp should match @Image 4. Use smooth rotating camera movements to display all angles. Lighting should be bright and clean to show all intricate details.
Video reference techniques
Replicating camera movements and cinematography
Video references excel at demonstrating specific camera techniques that are difficult to describe in text alone.
Place the character from @Image 1 in the corridor from @Image 2. Strictly follow all camera movement effects from @Video 1: tracking shot from behind as the character walks, camera circles around to the front with a low-angle perspective, then pans right 90 degrees to frame the doorway. Execute as a single continuous shot with no cuts.
Copying motion patterns and choreography
For dance, fight sequences, or specific movement patterns, video references provide frame-by-frame motion guidance.
Feature the martial artist from @Image 1 performing moves in the training hall from @Image 2. The character should execute the exact kick sequence shown in @Video 1: spinning back kick, transition to roundhouse kick, ending with an aerial spinning kick. Match the speed, height, and fluidity of the reference movements.
Replicating special effects and visual techniques
Video references can demonstrate particle effects, transitions, compositing techniques, and other visual effects for accurate reproduction.
The character from @Image 1 performs a magical transformation. Reference the particle effects from @Video 1: glowing particles rise from the ground, swirl around the character, brightness intensifies, then particles burst outward revealing the transformed appearance from @Image 2.
Audio reference techniques
Background Music Integration and Mood Setting
Audio references establish the emotional tone and pacing of your video through music selection.
Create a 15-second motivational fitness video featuring the athlete from @Image 1 in the gym setting from @Image 2. Use the energetic music from @Audio 1 to establish an inspiring, powerful mood. Camera movements should match the driving rhythm of the music with dynamic push-ins and motion.
Beat Synchronization for Visual Changes
Sync scene transitions, cuts, or visual changes to specific musical beats for polished, professional pacing.
The character from @Image 1 changes outfits with each musical beat from @Audio 1. First outfit from @Image 2, cut to second outfit from @Image 3 on the first beat, third outfit from @Image 4 on the second beat, fourth outfit from @Image 5 on the third beat. Each cut happens precisely on the beat. Use quick cuts with no transition effects.
Voice Timbre and Dialogue Matching
When specific voice characteristics matter, reference audio or video files containing the desired voice quality.
The narrator's voice should match the deep, authoritative timbre from @Audio 1. The narration text: "In a world transformed by technology, one person dares to question everything." Deliver with the same pacing and dramatic emphasis as the reference.
Complex multi-reference examples
Combining All Input Types for Commercial Production
Example: Product Commercial
Context: A minimalist studio with @Image 1 as the environment reference: white seamless background, dramatic side lighting. Reference: @Image 2 and @Image 3 show the product (wireless headphones) from the front and the side. @Video 1 for the camera movement, a slow rotating dolly. @Audio 1 for the background music. Shot 1: Wide shot establishing the headphones floating at the centre of the frame. The camera begins a slow circular dolly around them, matching the path in @Video 1. Shot 2: The headphones gently unfold as the camera continues around. <A soft mechanical click.> Shot 3: The camera pushes in to a close-up as the LEDs on the ear cups light up. <A soft power-on tone.> Throughout: (upbeat electronic music from @Audio 1). Style: high definition, clean product-commercial finish, crisp reflections, natural colour. Do not generate a logo or watermark. Keep it subtitle-free.
Multi-Character Scene with Dialogue
Example: Narrative Scene
Context: A corporate conference room in daytime, with the modern interior of @Image 1: large windows, a long table. Reference: Define the woman in the navy suit in @Image 2 as the CEO. Define the man in the grey suit in @Image 3 as the Investor. @Video 1 for the shot-reverse-shot camera rhythm between speakers. @Audio 1 for the tense ambient music. Shot 1: Wide shot, the CEO and the Investor at opposite ends of the table. The CEO stands, gesturing firmly, and says {This merger is our only option.} Shot 2: Cut to a medium shot of the Investor. He leans back, arms crossed, and replies flatly: {I've heard that before.} Shot 3: Cut to a medium shot of the CEO. She sits down heavily, jaw tight. Shot 4: Wide shot as the Investor stands. The camera tracks him smoothly as he walks to the window. <Footsteps on a hard floor.> Throughout: (tense ambient music from @Audio 1, low in the mix). Style: high definition, cool corporate tones, soft window light. Faces stay consistent with no deformation. Keep it subtitle-free.
Advanced Seedance 2.0 features
Video extension for continuous narratives
Seedance 2.0 can extend an existing video with new content that continues the story. It works in both directions: you can generate what happens after a clip, and you can generate what happened before it, which is useful when you have a strong shot but no way into it.
How Video Extension Works:
- Upload your existing video
- Say which direction you want, and describe what happens
- The model extracts the transition frames itself and generates the continuation
Refer to the clip as @Video 1, not "reference @Video 1". This is an editing task, and the word reference can push the model into generating a lookalike video instead of continuing yours.
Example: Extending a Coffee Shop Scene
Existing Video: 10-second clip of a person sitting at a cafe table, looking at a laptop
Generate the content after @Video 1. The person closes the laptop, picks up their coffee cup, and takes a sip while gazing out of the window, then sets the cup down and stands. The camera stays in a medium shot, holding the composition and lighting of the original.
The original footage is not regenerated, so the part you already like stays exactly as it is. By default the output carries only the tail of the source clip, not the whole thing. If you want the original included in the returned video, ask for it: "Extend @Video 1 backward, [description], and then end with @Video 1."
Two limitations to plan around
Extension is genuinely useful, but it is not invisible, and you will save yourself a bad afternoon by knowing this up front.
Joins can jump. At the point where the new segment meets the original, you may see a small jump or a frame of rollback. There is no prompt that reliably fixes this today. The practical repair is in your editor: trim about 6 frames from the end of the outgoing clip and 1 frame from the start of the incoming one, at every join. Better still, plan your clips so the join lands on a cut rather than mid-motion, where a discontinuity is invisible by design.
Quality decays if you chain extensions. Feeding a generated video back in as the source for another extension degrades the image, and the damage compounds with each pass. It shows up first as mottled colour blotches on faces. Keep the chain short and feed in the highest-quality source you have.
If you genuinely need to extend repeatedly, ByteDance documents an odd but effective workaround: convert the clip into a white-model video first, then continue from that. The prompt is roughly "Convert the video into a white 3D model. All characters become pure white 3D models, no colour, no texture, no shadows, pure white background, stable structure, smooth motion." Stripping the colour and texture gives the next pass a clean structural base to build on instead of compounding the artefacts of the last one.
Extension Best Practices:
- Keep extensions relatively short (5-8 seconds) for best continuity
- Clearly describe the connecting action between the original end and new content
- Mention elements that should remain consistent (camera angle, lighting, character position)
- If the original video has audio, reference that audio style for the extension
Video fusion and multi-clip transitions
Create seamless transitions between multiple existing video clips by generating bridging content.
Example: Connecting Two Locations
Existing Videos:
- @Video 1: Character walking in urban street (ends with character approaching corner)
- @Video 2: Same character entering apartment (starts with door opening)
Create a 5-second transition segment between @Video 1 and @Video 2. The character from the end of @Video 1 rounds the corner, walks up exterior apartment steps visible in background of @Video 2's opening frame, reaches the door, and begins opening it (connecting to @Video 2's start). Match the character's appearance, walking pace, and movement style from both reference videos. Lighting transitions from outdoor daylight at the start to the interior lighting of @Video 2 at the end.
This generates a bridge clip that smoothly connects two separate shoots, maintaining character and narrative continuity.
Character replacement in existing videos
Swap characters or subjects in videos while preserving all other elements including camera work, motion, and scene details.
Example: Music Performance Replacement
In @Video 1, replace the female lead singer with the male artist from @Image 1. The performance actions should exactly replicate those in the original video: microphone handling, body movements, facial expressions, and interaction with the band. The replacement artist should match the timing and energy of the original performance frame-by-frame. All other elements remain unchanged: band members, stage, lighting, camera movements.
Use Cases for Character Replacement:
- Testing different talent in commercial concepts
- Creating variations of the same scene with different actors
- Updating existing footage with new brand ambassadors
- Producing content for different regional markets with localized talent
Storyline subversion and narrative alteration
Completely change the narrative direction or outcome of existing video while maintaining the visual and technical elements.
Example: Relationship Drama Reversal
Original Video (@Video 1): Romantic scene where man proposes to woman on a bridge, she says yes, they embrace
Subvert the storyline of @Video 1. The scene begins identically: the man kneels and opens the ring box. However, the woman's expression shifts from surprised joy to shocked realization. She steps back, shaking her head. The man's face changes from hopeful to cold and calculating. He stands slowly, his demeanor becoming menacing rather than loving. The woman says "You were lying to me from the very beginning!" The man responds with an icy smile: "This is what you owe my family." The confrontational ending replaces the original romantic embrace. Maintain all camera angles and movements from @Video 1.
This technique allows complete narrative redirection while preserving the cinematography and production value of existing footage.
One-take continuous long shots
Create seamless long-take sequences that follow subjects through multiple environments without cuts.
Example: Urban Chase Sequence
@Image 1, @Image 2, @Image 3, @Image 4, and @Image 5 depict a one-take tracking shot following a runner. Sequence: Begin at street level (@Image 1) with a wide shot as the runner enters frame from the right, running at full speed. Camera picks up and follows from behind as runner reaches the building entrance (@Image 2). Continue tracking as runner bounds up the interior staircase (@Image 3), maintaining close following distance. Emerge onto the rooftop level (@Image 4), camera still tracking from behind. Runner reaches the roof edge. Camera moves around to the front of the runner for the final frame, then cranes up to overhead perspective showing city skyline (@Image 5). Camera: Continuous handheld-style tracking throughout. No cuts. Slight camera shake for urgency and realism. Smooth movement transitions between environments. The shot order above carries the pacing; do not pin the segments to fixed second counts.
Creative template replication
Copy the structure, style, and techniques from reference videos while substituting your own subjects and branding.
Example: Adapting Commercial Style
Reference: @Video 1 shows a high-end perfume commercial with specific camera techniques, transitions, and pacing
Create a luxury watch commercial by referencing the advertising style and structure of @Video 1. Use the same camera techniques: smooth dolly movements, dramatic lighting reveals, close-up detail focus, and elegant pacing. Replace the perfume bottle with the watch from @Image 1. Maintain the sophisticated color grading, transition timing, and rhythm from the reference. The environment should be minimalist and modern like @Image 2. Use the orchestral music from @Audio 1 to match the premium feel.
Seedance 2.0 use cases and examples
This section demonstrates Seedance 2.0 applications across different industries and complexity levels. Each industry includes basic, intermediate, and advanced examples showing progressive skill development.
Commercial and advertising production
Basic: Single Product Static Showcase
Scenario: Simple product display for e-commerce
Display the smartwatch from @Image 1 centered on the white background from @Image 2. Camera slowly rotates a full 360 degrees around the product, holding the same distance throughout. Lighting is clean and bright with no harsh shadows. Late in the shot the watch face illuminates, showing the time display. Use the subtle ambient electronic music from @Audio 1. Keep it subtitle-free. Do not generate a watermark.
Complexity Level: Single image reference, one camera move, one event
Intermediate: Multi-Angle Product Demo
Scenario: Tech product demonstration showing multiple features
Context: A clean studio, using @Image 1 as the lighting reference: soft even illumination against a minimal background. Reference: @Image 2 (front view of the wireless earbuds), @Image 3 (side view), @Image 4 (charging case, open). @Audio 1 for the background music. Shot 1: Overhead shot looking down at the open case. The lid closes on its own. <A satisfying click.> Shot 2: Front three-quarter angle. The case opens and the earbuds rise slightly out of it. <A soft whoosh.> Shot 3: The camera pushes in as a single earbud lifts free and rotates slowly to show every side. Shot 4: Close-up of the LED indicator on the case, pulsing to show it charging. Throughout: (upbeat music from @Audio 1). Style: high definition, clean commercial finish, natural colour. Keep it subtitle-free. Do not generate a logo or watermark.
Complexity Level: Multiple images, several shots, varied camera angles, audio sync
Advanced: Full Commercial with Scene Transitions
Scenario: 15-second lifestyle commercial showing product in use across multiple settings
Context: A lifestyle commercial for the wireless headphones in @Image 1 and @Image 2 (two angles). Reference: Define the young professional in @Image 3 as the Commuter (headshot). @Image 4 for the urban street, @Image 5 for the gym. @Video 1 for the dynamic camera style in the gym. @Audio 1 for the music that drives the whole piece. Shot 1: The camera tracks alongside the Commuter at medium distance as she walks a busy street wearing the headphones from @Image 1. <Street noise, dense and close.> She taps the ear cup, and the street noise drops away to nothing, leaving only the music. Shot 2: Cut on the beat. She is now at a desk on a video call, headphones on, in a bright home office. The camera pushes in to a close-up of the ear cup and its indicator light. <Soft keyboard typing under the music.> Shot 3: Cut on the beat to the gym from @Image 5. She trains hard and the headphones stay put. The camera moves with the energy, following the style of @Video 1, then pulls back to a wide shot. Throughout: (music from @Audio 1 carrying across all three locations). Style: 21:9, professional colour grading matching the cool modern tones of @Image 1, high definition. Cut on the beat. Do not generate a logo or watermark.
Complexity Level: 5 images, 1 video, 1 audio. Layered audio, three locations, professional cinematography.
This example sits at the top of the recommended asset range, and the home office is described in words rather than given its own reference image. That is a deliberate trade. A three-location commercial is exactly where the temptation to upload a picture of everything is strongest, and exactly where doing so starts to blur the subject.
Social media content creation
Basic: Trending Style Quick Cut Video
Scenario: Simple social media content with popular transition effect
Define the woman in @Image 1 as Subject 1. Subject 1 stands centred in frame against the bright background from @Image 2. Shot 1: She raises a hand in a quick gesture. On the gesture, jump cut to Subject 1 in the outfit from @Image 3, in the same position and pose. Shot 2: Another hand gesture, and jump cut to the outfit from @Image 4, same position. Shot 3: She holds the final pose and smiles. Throughout: (upbeat music from @Audio 1). Every cut lands on a beat. Her face stays identical across all three outfits. Keep it subtitle-free.
Complexity Level: Multiple image references, beat synchronization, simple transition effect
Intermediate: Multi-Location Story Sequence
Scenario: Day-in-the-life vlog style content
Context: A "day in the life" montage. Reference: Define the woman in @Image 1 as the Creator. @Image 2 (morning coffee shop), @Image 3 (co-working space), @Image 4 (outdoor park). @Video 1 for the handheld camera movement style. @Audio 1 for the background music. Shot 1: Coffee shop from @Image 2. The Creator enters, orders at the counter, and turns to the camera with a coffee in hand, handheld style from @Video 1. <Low cafe murmur under the music.> Shot 2: Co-working space from @Image 3. She works at a laptop, typing, then looks up at the camera and smiles. <Keyboard typing.> Shot 3: Close-up of the laptop screen. Shot 4: Park from @Image 4, golden hour. She sits on a bench, closes the laptop, stands, stretches her arms overhead, and walks toward the camera. <Birdsong and light wind.> Throughout: (upbeat vlog music from @Audio 1). Style: handheld vlog look following the movement of @Video 1, mixing medium shots and close-ups, cutting between locations on the beat. High definition, warm natural colour. Keep it subtitle-free.
Complexity Level: Multiple locations, handheld style reference, audio layering, personality-driven content
Advanced: Viral-Style Complex Visual Effects
Scenario: High-production social media content with trending effects
Context: A transformation video set against the urban background from @Image 4. Reference: Define the woman in @Image 1 as the Dancer. @Image 2 (starting outfit, casual streetwear), @Image 3 (ending outfit, performance costume), @Video 1 (choreography for the arm movement and the spin), @Video 2 (particle effect style), @Audio 1 (music). Shot 1: The camera circles the Dancer slowly. She stands still in the streetwear from @Image 2. Shot 2: She raises both arms in the movement from @Video 1. At the peak of the raise the image glitches with digital distortion. <A hard glitch sound.> Shot 3: Particle effects matching @Video 2 burst up from the ground and swirl around her. The camera continues its rotation. <A rising whoosh.> Shot 4: A flash of light. As it fades she is in the performance costume from @Image 3, mid-spin, following the choreography in @Video 1. Shot 5: She completes the spin and lands in a held pose. The camera settles front-on. The environment behind her is now the stage from @Image 5. <An impact on the landing.> Throughout: (high-energy music from @Audio 1, building to a climax). Style: high contrast, saturated colour, beat-synced effects, high definition. Her face stays identical before and after the transformation. Keep it subtitle-free.
Complexity Level: Multiple complex references, precise choreography matching, special effects replication, advanced audio sync, trending style integration
Film and entertainment production
Basic: Atmospheric Establishing Shot
Scenario: Scene-setting shot for narrative content
Cinematic establishing shot of the abandoned mansion from @Image 1 at night. Single shot: the camera starts wide on the full building and its overgrown grounds, then pushes in slowly toward the main entrance. One continuous move, no cuts. Dark, moody atmosphere, moonlight breaking through cloud. Every window is dark except one on the second floor, where a faint light flickers. <Wind moving through trees.> Throughout: (ominous ambient tone from @Audio 1). Style: filmic, high definition, deep shadow, cold moonlight. Keep it subtitle-free.
Complexity Level: Single image, basic camera movement, atmosphere building
Intermediate: Dialogue Scene with Shot Reverse Shot
Scenario: Two-character conversation with professional coverage
Context: An interrogation room with the stark environment of @Image 1: one overhead light, a metal table, two chairs. Reference: Define the stern middle-aged man in @Image 2 as the Detective. Define the young man in @Image 3 as the Suspect. @Video 1 for the shot-reverse-shot camera rhythm. @Audio 1 for the tense ambient music. Shot 1: Wide shot of both men at the table. The Detective leans forward, hands clasped. The Suspect will not meet his eye, fingers tapping the tabletop. <A metal chair creaks.> Shot 2: Cut to a medium close-up of the Detective, slightly low angle. Unblinking, he says {We know you were there that night.} Shot 3: Cut to a medium close-up of the Suspect, slightly high angle. His eyes dart away, then he composes himself: {I don't know what you're talking about.} Shot 4: Cut back to the wide shot. The Detective slides a photograph across the table. The Suspect's eyes widen. The Detective leans back. Throughout: (tense music from @Audio 1, low) and room tone. <The photo slides across metal.> Style: harsh dramatic light, high definition, filmic. Faces stay consistent, no deformation. Keep it subtitle-free.
Complexity Level: Two character images, specific camera technique reference, dialogue pacing, psychological tension
Advanced: Action Sequence with Complex Choreography
Scenario: Fight scene with specific martial arts choreography
Context: A rooftop at sunset, environment from @Image 1: HVAC units, distant skyline, hard orange sky. Reference: Define the fighter in @Image 2 as the Hero (headshot), with costume from @Image 3. Define the three opponents in @Image 4 as the Opponents. @Video 1 for the fight choreography: the dodge, the spinning kick, the grapple. @Video 2 for the camera style: circling the action, cutting on impact. @Audio 1 for the music. Shot 1: Wide shot. The Opponents surround the Hero. The camera circles the group slowly. Wind moves their clothing. Shot 2: Close-up of the Hero's face. Jaw set, breathing steady. Shot 3: The first opponent charges. The Hero dodges right, matching the movement in @Video 1. Shot 4: The Hero executes the spinning kick from @Video 1 and connects. The camera follows the kick in a medium shot. <A heavy impact.> Shot 5: The second opponent closes in. The Hero grapples: grab, pivot, throw, as in @Video 1. The camera circles the action in the style of @Video 2. Shot 6: The third opponent swings. The Hero ducks underneath in slow motion. Shot 7: The Hero stands, the Opponents on the ground around him. The camera pushes in to a close-up, sunset behind him throwing him into silhouette. Throughout: (music from @Audio 1, building to a climax), <wind across the rooftop, cloth movement, heavy breathing>. Style: warm sunset tones, high contrast, filmic, high definition. Bodies stay anatomically stable, with no deformed limbs. Keep it subtitle-free.
Complexity Level: Four image references, two video references (choreography and camera style), audio reference, complex action choreography, slow motion.
Set your expectations honestly here. Fast, high-impact action is the hardest thing to ask of this model, and it is where limbs deform and physics gives way. That is exactly why this prompt leans on a choreography reference video rather than trying to describe the kick in words: showing the motion is far more reliable than describing it. Even so, expect to generate this one several times. If a shot keeps failing, the fix is usually to slow the action down or break it into two shots, not to add more adjectives.
Professional workflow applications
Video Extension for Project Continuity
Scenario: Extending previously shot footage with additional content
Existing Video: 8-second shot of CEO walking through modern office, ending at conference room door
Generate the content after @Video 1. The CEO from the end of the video opens the conference room door and enters. Inside, the conference room matches the design from @Image 1: large table, floor-to-ceiling windows with city view. Three executives from @Image 2, @Image 3, and @Image 4 are already seated and look up as CEO enters. CEO walks to the head of the table and sits down. Camera follows CEO through doorway with smooth tracking shot, then cuts to wide shot showing full conference room once CEO is seated. Maintain the same professional color grading and lighting style from @Video 1.
Use Case: Adding to existing professional video assets without reshoots
Template-Based Bulk Content Creation
Scenario: Creating multiple social media videos with consistent style
Master Template Prompt (Video 1):
Product showcase for [the product in @Image 1] against the white background from @Image 2. Shot 1: The camera rotates a full 360 degrees around the product. Shot 2: A graphic callout appears beside the product highlighting its key feature. Shot 3: The logo from @Image 3 resolves in the centre of frame. Throughout: (music from @Audio 1). Style: clean commercial finish, high definition, soft even light. Do not generate a watermark.
Variation Prompts: Replace @Image 1 with different products while maintaining @Image 2, @Image 3, and @Audio 1 for brand consistency
Use Case: Scalable content production for product catalogs, maintaining brand identity across multiple assets
Multi-Language Adaptation
Scenario: Creating regional variations of the same commercial
Base Prompt:
30-second commercial structure from @Video 1. Replace narration with [Language] voice matching @Audio 1's tone and pacing. Character from @Image 1 remains the same. Text overlays change to [Language] versions matching timing from @Video 1.
Use Case: International marketing campaigns requiring localized versions with consistent visual branding
Best practices for Seedance 2.0
The CRAFTS prompting framework (detailed)
Professional results in Seedance 2.0 require structured prompt engineering. The CRAFTS framework provides a systematic approach that ensures all critical elements are specified:
C - Context: Establish Scene and Environment
Define where and when the action takes place. This includes:
- Physical location and setting
- Time of day or historical period
- Atmospheric conditions (weather, lighting quality)
- Overall mood and tone
- Environmental details that matter to the story
Example: "In an underground nightclub at 2 AM, with the moody atmosphere from @Image 1. Hazy air from smoke machines, a low ceiling lit by the glow of the bar, packed dance floor in the background."
R - Reference: Define Subjects, Then Specify @ Mentions and Exact Purpose
This is where multimodal power lives. Name the subject first, then be explicit about what each reference contributes:
- Define each subject by 2 to 3 stable features, and reuse that label every time
- State the @ mention clearly
- Specify exactly what aspect of that reference to use
- Clarify what NOT to use if the reference contains multiple elements
- Lead with the asset that has to be most accurate
Example: "Define the woman with the shaved head and leather jacket in @Image 1 as Ada. @Image 1 for Ada's facial features and hair only, not the clothing. @Image 2 for the leather jacket costume. @Video 1 for the walking pace and confident stride. @Audio 1 for the electronic background music."
A - Action: Describe Character and Object Movements
Detail what happens in the scene: the verbs of your video:
- Character movements and gestures
- Object interactions (picking up, setting down, throwing)
- Facial expressions and emotional reactions
- Interactions between multiple subjects
- Physics-based events (things falling, liquids pouring, smoke rising)
Example: "Character enters from frame left, walking with the confident stride from @Video 1. Eyes scan the crowd briefly, then lock onto someone off-screen. Slight smile forms. Character adjusts jacket collar with right hand, then begins moving forward through the crowd with purpose."
F - Framing: Define Camera Work and Cinematography
Use proper cinematography terminology to specify shot composition:
- Shot types: Wide shot, medium shot, close-up, extreme close-up, over-the-shoulder, point-of-view
- Camera movements: Dolly in/out, tracking shot, pan left/right, tilt up/down, crane up/down, handheld, steadicam
- Angles: Low angle, high angle, eye level, dutch angle
- Special techniques: Hitchcock zoom, whip pan, rack focus, shallow depth of field
Example: "Open with wide shot establishing the full nightclub environment. As character enters, camera picks up and begins tracking alongside in medium shot. When character stops to scan crowd, push in slowly to medium close-up. Cut to character's POV shot looking through crowd. Cut back to close-up of character's face as smile forms. Resume tracking shot as character moves through crowd, camera following from behind."
T - Timeline: Order Your Shots and Let the Model Find the Pacing
Break your sequence into numbered shots, in the order events happen:
- Label them Shot 1, Shot 2, Shot 3, most important first
- Give each shot one camera move, one beat of action, one place
- Put the sound for that shot inside that shot
- Do not assign each shot a duration
Precise timing is the model's weakest control surface. Pinning segments to exact second counts can produce worse output than leaving them alone, so describe the order of events and let the model breathe.
Example: "Shot 1: Wide shot of the club. Ada enters and starts walking. Shot 2: The camera tracks alongside her in a medium shot as she scans the crowd. Shot 3: Close-up as a small smile forms. Shot 4: Her POV, looking through the crowd. Shot 5: The camera resumes tracking from behind as she moves forward. Throughout: (electronic music from @Audio 1 at moderate volume)."
S - Style and Constraints: Lock the Look and Rule Out the Artefacts
Close every prompt with the finish and the guardrails:
- Image quality: "high definition, rich detail, filmic texture, soft light"
- Style: name it explicitly, even if a reference image implies it
- Constraints: "Keep it subtitle-free." "Do not generate a watermark." "Do not generate a logo."
- Stability: "The character's face stays consistent, with no deformation. Motion is smooth, with no stutter or flicker."
Complete CRAFTS Example: Corporate Training Video
Context: A modern conference room in the morning, natural window light from frame right. Environment matches the interior of @Image 1: glass walls, contemporary furniture, screens on the wall. Reference: Define the woman in the navy blazer in @Image 2 as the Trainer. @Image 3 for the group of trainees seated around the table. @Video 1 for the Trainer's hand gestures and body language while explaining. Shot 1: Wide shot of the whole room from the corner, establishing the Trainer at the head of the table and the trainees around it. Shot 2: Medium shot of the Trainer from a front three-quarter angle. She gestures toward the screen, then turns to the group with an open smile. Shot 3: Over-the-shoulder from behind the Trainer, showing the trainees listening. One leans forward. Shot 4: The camera tracks alongside the Trainer as she walks the length of the table, making eye contact as she goes. Shot 5: Close-up of a trainee taking notes. <Quiet keyboard tapping.> Shot 6: Medium shot of the Trainer back at the head of the table, closing with a confident gesture. Throughout: (corporate background music from @Audio 1, quiet, dipping under the dialogue). Style: high-definition corporate documentary look, warm neutral tones, soft light. Faces stay consistent with no deformation, motion is smooth. Keep it subtitle-free. Do not generate a logo or watermark.
Input preparation strategy
Image Reference Optimization
Quality input creates quality output. Prepare image references strategically:
For Character Consistency:
- Use a face-only headshot, shot straight-on, neutral expression, minimal background
- Add one full-body shot for costume and proportion. That is the whole set
- Do not add more angles. A front/side/three-quarter turnaround makes identity drift worse, not better, because the model reads the angles as different people
- Avoid heavy filters or effects that might confuse the model
- If the costume matters, let the full-body shot carry it rather than adding detail crops
For Style and Aesthetic:
- Select images that clearly demonstrate the desired visual treatment
- Ensure color grading is consistent with final vision
- Include images showing the specific lighting approach you want
- Consider texture and detail level: high detail references produce high detail outputs
For Products and Objects:
- Photograph against simple backgrounds for focus
- Show multiple angles to ensure accurate reproduction
- Include close-ups of important details (logos, textures, specific features)
- Ensure lighting shows form and dimension clearly
Video Reference Optimization
For Camera Movement:
- Trim videos to show only the specific camera move you want to replicate
- Ensure the movement is clearly visible and not obscured by action
- Shorter clips (3-5 seconds) focused on one technique work better than longer clips with multiple techniques
- Use highest quality video available: compression artifacts affect understanding
For Motion and Choreography:
- The action should be clearly visible without obstruction
- Ensure lighting adequately shows body position and movement
- Multiple angles of the same action can help if available
- Consider slowing down fast movements when creating reference clips
For Special Effects:
- Isolate the specific effect you want to replicate
- Ensure effect is clearly visible against background
- If effect has specific timing, include that timing in reference
Audio Reference Optimization
For Music and Rhythm:
- Use high-quality audio files (avoid low-bitrate compressed audio)
- Trim audio to the section with the most relevant rhythm or mood
- Ensure audio clearly demonstrates what you want (beat, pace, mood)
- Consider starting audio at a strong beat for easier synchronization
For Voice and Dialogue:
- Use clear recordings with minimal background noise
- Ensure the specific vocal characteristic you want is prominent
- Keep reference clips short and focused on the relevant vocal quality
Choosing your four or five assets
There is no combined file cap, so nothing stops you uploading nine images. The discipline has to come from you. Use this to decide what earns a slot.
The four functional roles. A complete brief usually needs one of each, and rarely more:
- Character anchoring (1-2 images): locks who the subject is
- Scene tone-setting (1 image): locks the environment and style
- Camera movement (1 video): locks the shot language and rhythm
- Rhythm and atmosphere (1 audio): locks emotion and timbre
Four questions to cut the rest:
- Will removing this reference actually change the result? If not, cut it.
- Can this be said in text instead? Environments, moods, and colour grades usually can. Faces, exact camera moves, and specific choreography usually cannot. Upload what is hard to describe. Describe everything else.
- Is another asset already doing this job? Two references competing for the same role is worse than one.
- Is this a must-have or a nice-to-have? Cut the nice-to-haves first.
Example decision process
You are making a music video and have fifteen candidate references:
- 4 images: the artist from different angles
- 3 images: the performance venue
- 2 images: lighting setups
- 2 videos: dance choreography and camera movement
- 2 audio files: music track and ambient sound
- 2 images: costume details
Applying the framework:
- Keep (character anchoring): 1 headshot of the artist, plus 1 full-body shot. Not the four-angle set. A multi-angle sheet is the single most reliable way to make the model think it is looking at several different people
- Keep (scene): 1 venue image, the most representative
- Keep (camera): 1 video. Choose between the choreography and the camera-movement clip based on which one you could not write down in words
- Keep (rhythm): the music track
- Describe in text: both lighting setups, the costume details, the ambient audio, and the other two venue images
Result: 5 assets. The other ten pieces of information still reach the model, they just arrive as words instead of files, which is exactly where they belong.
Consistency techniques for multi-shot projects
Character Consistency Across Generations
Maintaining the same character appearance across multiple video generations requires systematic reference management:
Use a headshot plus a full-body shot. Nothing more.
This is the single most counter-intuitive rule in the model, and getting it wrong causes the exact problem you were trying to solve.
The instinct is to give the model a character turnaround: front, side, three-quarter, back. Do not. Seedance 2.0 frequently reads a multi-view sheet as several different people. That makes identity drift worse, not better, and it is a leading cause of the "twins" bug, where two identical copies of your character show up in the same frame.
What works is two images with clearly separated jobs:
- A face-only headshot. Head only, neutral expression, minimal shoulders, neck, or background. This is the image the model leans on for identity, so the face needs to dominate the frame. If the face is a small region inside a busy full-body shot, the model under-weights it and the surrounding clutter bleeds in.
- A full-body shot for styling, costume, and proportion.
Then split the roles explicitly in the prompt, and put the face first:
Define the woman in @Image 1 as Subject 1. Subject 1's facial features reference @Image 1 (the headshot). Her outfit and styling reference @Image 2 (the full-body photo).
Note that this applies to faces. For products and objects, multiple angles are genuinely useful and you should supply them.
The mixed-reference mistake: do not hand the model one composite image that crams the face, the pose, the outfit, and a detail crop into a single picture. The face ends up occupying a small fraction of the pixels, gets a correspondingly small share of the model's attention, and identity drifts.
If your character drifts anyway, there is a consequence worth knowing about: a drifting face can wander toward resembling a real public figure, and that can get the generation blocked at review even though you never asked for a celebrity. So identity drift is not only a quality problem. It is one of the quieter reasons a clean-looking prompt comes back rejected. Our guide to Seedance 2.0 prompts getting flagged covers the rejection side in detail.
Keeping a consistent label Use the same name for the character in every prompt, and specify "maintaining exact appearance from @Image [X]".
Feature the detective from @Image 1 (maintain exact facial features, hairstyle, and clothing from this reference). In this scene, the detective enters the warehouse from @Image 2. All physical characteristics of the detective must match @Image 1 precisely: same face, same coat, same build.
Style Consistency Across Scenes
For projects requiring multiple shots with consistent visual treatment:
Technique 1: Style Reference Template Select one image that perfectly captures your desired visual style:
- Color grading
- Lighting approach
- Composition style
- Texture and detail level
Include this same style reference in every generation prompt:
Maintain the visual style from @Image 1 throughout: moody blue color grading, high contrast lighting, film grain texture, shallow depth of field.
Technique 2: Previous Output as Reference Use earlier successful generations as references for later shots:
Create the next scene maintaining the exact visual style from @Video 1 (my previous generation). Color grading, lighting approach, and overall aesthetic should match precisely.
Temporal Continuity for Sequential Shots
When creating shots that connect sequentially:
Technique 1: Overlap Description Describe how the new shot connects to the previous:
This shot picks up exactly where @Video 1 ended. The character who was facing the door at the end of @Video 1 now turns toward camera and begins speaking. Position and lighting should match the final frame of @Video 1.
Technique 2: Transition Specification Clearly state the connection point:
Start this generation with the same camera angle and position where @Video 1 concluded. The character should be in the same position, mid-gesture, and this shot continues the motion smoothly.
Common pitfalls to avoid
Pitfall 1: Vague Reference Usage
Problem: "@Image 1 as reference" without specifying what aspect to reference
Solution: Always state exactly what the reference provides: "@Image 1 for character's facial features and expression, not the background or lighting"
Pitfall 2: Contradictory Instructions
Problem: "Fast-paced action scene with slow, contemplative camera movements and calm ambient music"
Solution: Align all elements: action pace, camera energy, music tempo, editing rhythm: toward a consistent goal
Pitfall 3: Over-Complicating Prompts
Problem: Uploading a dozen barely-differentiated references and writing a 500-word prompt full of conflicting details. ByteDance's guidance is blunt about this: do not hand the model your whole script, and do not use the full asset allowance
Solution: Use fewer, higher-impact references with clear, structured prompts following the CRAFTS framework
Pitfall 4: Ignoring Duration Limitations
Problem: Trying to fit 30 seconds of detailed action into 15-second generation
Solution: Break complex sequences into multiple generations or simplify action to fit time constraints
Pitfall 5: Under-Specifying Camera Work
Problem: "Camera moves around" without specific direction
Solution: Use precise cinematography terms: "Camera dollies in from wide shot to medium close-up over 5 seconds, maintaining eye-level perspective"
Pitfall 6: Neglecting Audio Integration
Problem: Treating audio as afterthought or only mentioning "add music"
Solution: State what the audio is for, and put the sound inside the shot it belongs to: "(driving rhythm from @Audio 1 throughout). Every cut lands on a beat."
Pitfall 7: Inconsistent Reference Quality
Problem: Mixing high-resolution professional photos with low-quality compressed images
Solution: Maintain consistent quality across all references: don't let one poor-quality reference compromise the generation
Pitfall 8: Assuming Model Inference
Problem: "Make it look good" or "you know what I mean"
Solution: Be explicit about every important detail: the model executes your instructions, it doesn't interpret vague intent
Quick Troubleshooting Guide
Issue: Character appearance changes between generations Solution: Use identical character reference image in each prompt, explicitly state "maintain exact appearance from @Image X"
Issue: Camera movement isn't matching reference Solution: Add more specific description of the camera movement in text, break complex movements into stages
Issue: Style doesn't match reference Solution: Describe the specific style elements in text alongside the reference: "Match @Image 1's color grading: desaturated blues, high contrast, crushed blacks"
Issue: Timing feels off Solution: Resist the urge to add timecodes, which tends to make things worse. Reduce the number of shots you are asking for in one clip, and give the important beat its own shot.
Issue: Audio doesn't match mood Solution: Describe the audio's role explicitly rather than pointing at the file: not "@Audio 1" but "(tense, building suspense from @Audio 1)".
Issue: The voice does not sound like my reference audio Solution: Audio references alone under-deliver on timbre. Describe the voice in words as well as referencing it: "Use the low, warm, slightly grainy middle-aged male voice of @Audio 1 to say...". Keeping the written dialogue close in tone to the reference recording also helps.
Issue: Two identical copies of my character appear in one frame Solution: The "twins" bug. Stop using multi-view character sheets, use single-person reference photos, name each character and pair them with their image ("Ada (@Image 1) hands the file to Ben (@Image 2)"), and add a closing constraint: "Throughout the video, do not generate duplicate characters or a twin effect. Keep only one of each character in frame."
Issue: Subtitles or a watermark appear that I never asked for Solution: Add the constraint lines ("Keep it subtitle-free", "Do not generate a watermark"), strip any text out of your reference assets before uploading, and generate in landscape if you can, since spurious subtitles are markedly less common there than in portrait.
Issue: My animation style drifted back to live action Solution: A realistic reference image will drag the output toward realism unless you say otherwise. Name the style explicitly in the prompt ("2D Japanese animation style"), or convert the reference image into the target style before you generate from it.
Issue: A described effect comes out garbled, like a countdown that scrambles its digits Solution: Stop describing the effect and show it. Feed in a reference video of the effect and point at it: "the way the number appears should follow @Video 1." Motion logic is one of the things the model reads far better from footage than from text.
Issue: The video ends with a click or an abrupt cut-off noise Solution: A known artefact on clips with narration. Regenerate, or fix it in your editor by applying a short volume fade to the end of the audio track. It is a 20-second fix and not worth burning a generation on.
Issue: The image jumps or stretches, and my reference photo looks distorted Solution: Usually an aspect-ratio mismatch. If your reference image's ratio does not match the output ratio you selected, the model has to force a fit. Crop the image to match your target ratio before uploading, or set the output to the adaptive ratio so it follows the reference instead. Reference images also need to sit within a 0.4 to 2.5 aspect ratio and 300 to 6000px per side.
Conclusion
Seedance 2.0 represents a fundamental advancement in AI video generation through its comprehensive multimodal approach. By accepting images, videos, audio, and text as inputs, it provides professionals with unprecedented control over the creative process: moving beyond text-only prompts to true show-and-tell direction.
Seedance 2.0's position in the AI video landscape
The multimodal capability distinguishes Seedance 2.0 from competing platforms. While Kling, Veo, and Sora offer impressive text-to-video capabilities, Seedance's integration of direct video and audio references enables precise reproduction of camera work, motion patterns, and rhythm synchronization that would be difficult or impossible to achieve through text description alone. This positions Seedance as the tool of choice for professionals who need exacting control over visual style, character consistency, and cinematic execution.
The platform continues to evolve with regular capability enhancements and expanded feature support. Mastering the multimodal reference system and CRAFTS prompting framework provides a foundation for increasingly sophisticated video creation as the platform develops.
Key takeaways
Multimodal Control: Seedance 2.0's combination of image, video, audio, and text inputs enables showing the AI exactly what you want rather than attempting to describe it entirely in words. This fundamental approach shift makes previously difficult specifications: exact camera movements, specific choreography, beat-synchronized editing: straightforward to achieve.
Strategic Comparison Advantages: Compared to Kling, Veo, and Sora, Seedance 2.0 offers unique capabilities in audio integration and video reference depth. The direct audio file upload and reference system enables precise mood control and beat synchronization. The video reference capability extends beyond style transfer to full motion and camera replication.
CRAFTS Professional Framework: The six-part CRAFTS methodology (Context, Reference, Action, Framing, Timeline, Style and constraints) covers everything a complete prompt has to specify. The two parts people most often skip are the ones that decide whether the generation works: defining the subject before referencing it, and closing the prompt with the constraints that rule out artefacts.
Sequence your shots, do not time them: the instinct to write second-by-second timecodes is the most common way to make Seedance output worse. Order your shots and let the model find the pacing.
Restraint beats volume: four or five well-chosen assets beat a dozen. Every additional reference makes it harder for the model to work out which features matter, and character turnarounds in particular actively cause identity drift.
Available on Morphic: Professional creators can access the full Seedance 2.0 family through Morphic, including Fast and Mini for cheap iteration and full Seedance 2.0 for 4K finals.
FAQs
Use the same character reference image in every generation where that character appears. In your prompt, explicitly state "maintain exact appearance from @Image X" and describe any variations (different clothing, expression) while emphasizing that facial features, build, and other identifying characteristics remain identical. For best results, use a clear, well-lit frontal photo as your master character reference.
Upload the video showing the desired camera work and reference it specifically: "@Video 1 for camera movement only." In your text prompt, describe the movement using cinematography terminology (dolly in, tracking shot, crane up). For complex movements, break them into separate shots rather than separate timecodes: "Shot 1: dolly in from wide to medium. Shot 2: pan right, holding distance." Keep each shot to a single camera move, since stacking a push, an orbit, and a pan into one shot is the fastest way to destabilise the image.
Upload your music track, then describe the sync in terms of beats rather than seconds, since the model's handling of precise timecodes is unstable. Something like: "(upbeat music from @Audio 1). Every cut lands on a beat," with each cut written as its own shot. Beat-relative instructions land far more reliably than "at the 6-second mark."
Use the video extension feature or fusion technique. For extension: upload your existing video and say "Generate the content after @Video 1" and describe the connecting action, setting the length in your generation settings rather than in the prompt. For fusion: create a bridging segment that references the ending of one clip and the beginning of another, explicitly describing the transition action that connects them.
Less directly than you would like, and this is the model's weakest control surface. Precise timecodes ("0-3 seconds: [action]") are unstable, and forcing exact durations onto segments can produce worse output rather than tighter output. Sequence your shots instead ("Shot 1: ... Shot 2: ...") and let the model find the pacing. If an action feels rushed, the reliable fixes are to ask for fewer shots in the clip, give that action its own shot, or raise the overall generation duration, rather than writing a tighter timecode.
Fewer than you are allowed. ByteDance recommends 4 to 5 assets: 1 to 2 character images, 1 scene image, 1 camera-movement video, 1 audio clip. Past that the model struggles to work out which features take priority, and you get style conflicts, blurry subject identification, and output that drifts from the brief. The rule of thumb: upload what is genuinely hard to describe (a specific face, an exact camera move, particular choreography) and put everything else in the text prompt. Note also that stability drops sharply once more than four reference people are involved.
Upload the video with the desired effect and specify: "@Video 1 for the particle effect technique only." In your text prompt, describe the effect in detail: when it occurs, how it moves, its visual characteristics. For best results, use reference clips where the effect is clearly visible and isolated: "Reference the glowing particle swirl from @Video 1 that rises from ground level and disperses as it reaches head height."
Upload an audio or video reference containing the desired voice and specify: "@Audio 1 for voice timbre and delivery style." In your prompt, describe the vocal characteristics: "The character speaks with the deep, authoritative tone from @Audio 1, delivering the line: [your dialogue text]."
Maintain consistent reference materials across all generations in your sequence. Use the same style reference image, the same character references, and similar prompts with only necessary variations. Include references to previous successful outputs: "Maintain the visual style from @Video 1 (previous generation)" to ensure continuity.
Use video extension. Generate your first segment, then upload it and say "Generate the content after @Video 1", describing what happens next. You can chain extensions, but keep the chain short: each pass degrades image quality, and the damage compounds, showing up first as mottled colour on faces. Expect a visible jump at the joins too, which you fix in your editor by trimming roughly six frames off the end of the outgoing clip and one off the start of the incoming one. For longer or fast-paced pieces, generating separate clips and cutting them together usually beats one long chain of extensions.
Seedance 2.0's primary differentiator is comprehensive multimodal input including direct audio file upload and the depth of its video reference and editing system. Both models take multimodal input, so the difference is less about what you can attach and more about what the model does with an uploaded clip: Seedance will edit inside it, extend it in either direction, and stitch clips together, not just imitate its style. That matters most on projects where footage already exists and has been signed off.
Seedance 2.0 accepts direct audio file uploads, and pulls three distinct things out of them: timbre, melody, and dialogue content. That means you can hand it the actual music track, or a recording of the voice you want, rather than describing either in words. Veo and Sora generate audio from text descriptions instead of accepting reference audio. One honest caveat: an audio reference alone tends to under-deliver on voice. Describe the timbre in text as well, for example "use the low, warm, slightly grainy middle-aged male voice of @Audio 1".
Seedance 2.0 and Kling 3.0 both generate up to 15 seconds in a single pass. Sora goes longer, up to 60 seconds. In practice the ceiling matters less than it sounds, because most commercial and social work is cut together from several short, high-quality clips rather than shot as one long take. Where Seedance pulls ahead is what happens after the clip exists: you can extend it forward or backward, edit inside it, and stitch clips together, which is a different kind of length than a longer single generation.
Seedance 2.0's multimodal approach provides more direct control for style replication because you can upload multiple reference images, video clips showing the style in motion, and audio that establishes mood. Rather than describing a style in text, you show examples from multiple angles. This typically results in more faithful reproduction of complex styles compared to text-only approaches.
Seedance 2.0's image reference system, when used correctly with consistent character images across prompts, provides strong character consistency. This capability is comparable to Kling's character consistency features but more controllable than Veo or Sora's text-based character descriptions. The key is using high-quality character reference images and explicitly stating "maintain exact appearance from @Image X" in each generation.
For commercial work the deciding factor is usually control rather than raw generation quality, and that is where Seedance 2.0's reference system earns its place: you can upload the exact product shots, the exact music track, and a clip of the exact camera move, which matters when a client has already signed off on all three. Seedance also lets you edit and extend existing footage, so approved material can be reused rather than regenerated. Veo is strong on prompt adherence and long-form extension, and remains a reasonable choice where the brief is looser and the length matters more than exact brand alignment.
They are built for different jobs. Sora generates longer single takes, up to 60 seconds, and is strong at inferring physics and complex scenes from text alone. Seedance 2.0 caps at 15 seconds per generation but gives you control Sora does not: direct audio upload, video references for motion and camera replication, several visual references at once, and the ability to edit and extend footage you already have. If your prompt is a paragraph of description, Sora is compelling. If you have assets and a specific result in mind, Seedance's reference system is the more controllable tool.
Both platforms offer motion reference capabilities, but Seedance 2.0's video reference system goes deeper. Kling provides motion brush and basic motion transfer, while Seedance allows uploading complete video clips and replicating not just motion paths but also camera work, editing rhythm, and complex choreography frame-by-frame. You can show Seedance an entire fight sequence or dance routine and have it replicate the motion precisely rather than describing it or drawing motion paths.
They share the same feature set and differ on quality and cost. Seedance 2.0 gives the highest quality and is the only one of the three that outputs 4K. Seedance 2.0 Fast trades some quality for speed and cost. Seedance 2.0 Mini is the cheapest and best suited to high-volume iteration. Fast and Mini both top out at 720p. A practical workflow is to iterate on Mini or Fast until the prompt and references behave, then run the final on full Seedance 2.0.
Images as JPG or PNG, video as MP4 or MOV, and audio as MP3. Output is MP4, with 4K encoded as 10-bit H.265, which some browsers and players will not preview directly even though the file is sound. Use the highest-quality sources you have: compression artefacts in a reference clip measurably degrade how well the model reads the motion in it.
Up to 9 images, 3 video clips, and 3 audio files. Video clips must be 2 to 15 seconds each and 15 seconds combined; audio files the same. There is no combined cap across types, so the maximum legal request is 15 files, but you should be nowhere near it: 4 to 5 assets is the recommended working range. One combination simply does not work: text plus audio with no visual, and audio on its own, will not generate.
Seedance 2.0 generates videos between 4 and 15 seconds in a single generation. You can select the specific duration in 1-second increments. For longer content, use the video extension feature to chain multiple generations or generate separate segments that can be edited together in post-production.
Yes, Seedance 2.0 through Morphic can be used for commercial production. Specific licensing and usage rights are governed by Morphic's terms of service. Review those terms for details on commercial use, client work, and any attribution requirements.
Within a single generation, yes. Quality holds across the clip. The place it does not hold is across chained video extensions: each time you feed a generated video back in to extend it, the image degrades a little, and it compounds. It shows up first as mottled colour blotches on faces. Keep extension chains short.
Yes. Seedance 2.0 supports 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16, set in your generation settings. There is no 2.35:1 option, so if you want a scope look, choose 21:9. One practical note: stray subtitles appear noticeably less often in landscape than in portrait, so if a piece is destined for vertical, it can be worth generating landscape and cropping down.
Seedance 2.0 is accessible through Morphic. Visit Morphic, create an account or log in, and access Seedance 2.0 through their video generation interface. The multimodal input system and @ reference functionality are integrated into Morphic's workflow.
Yes, you can use generated videos in several ways: as references for new generations (to modify specific elements), as inputs for video extension (to add continuation), in video fusion workflows (to connect with other clips), or export them for traditional video editing in standard editing software. Generated videos are yours to edit, combine, and refine through whatever workflow serves your project.