Wan 3.0 features and capabilities
Wan 3.0 is the next model expected in Alibaba Tongyi Lab's Wan video line. Early demos point to native 1080p output at roughly 30 seconds per clip, on top of the reference control and in-pass audio the line already runs.
| Capability | What it does | Best for |
|---|---|---|
| Long single takes | Holds one continuous shot for roughly 30 seconds | Ad spots, short scenes, long reveals |
| Native 1080p | Renders at full HD without an upscaling pass | Delivery-ready footage, large screens |
| Audio in the same pass | Returns a scored clip rather than a silent one | Sound-on scenes, dialogue, music beats |
| Reference-to-video | Locks a character, product, or set across shots | Series, multi-shot sequences, brand work |
| Prompt planning | Reasons through composition and motion before rendering | Complex briefs with several moving parts |
Long single takes
Length is the headline. Roughly 30 seconds in one continuous shot is about double what the line does today, and it changes the unit of work: a full ad spot or a short scene fits inside one generation instead of being cut together from several. You write it as one evolving motion, describing how the subject and camera move across the whole take rather than a single frozen frame.
Reference-led continuity
Reference-to-video is what keeps a sequence recognizable. Supply an image or clip of a character, product, or set, and the model carries that identity through the shot and from cut to cut. Naming what each reference is for matters as much as supplying it, so the model knows which one drives the subject and which fixes the location.
Audio in the same pass
Sound generates alongside the picture rather than being scored afterwards, so a scene can come back with room tone, effects, or dialogue already in place. That makes pacing part of the prompt: say where a beat lands and the picture can be cut to it.
Prompt planning
The Wan line plans before it renders, reasoning through composition and motion ahead of the first frame. A longer, better-organized prompt therefore pays off more here than on models that render straight from the text, which is what the prompt section below is built around.
On open weights
Alibaba open-sourced the earlier Wan families under Apache 2.0, and those weights are still the ones you can download and self-host. The generations since have been commercial API models. If self-hosting is a requirement for your pipeline, that is the detail to watch rather than any single specification.
How to get the best out of Wan 3.0
Each of Wan 3.0's strengths asks for something specific from the prompt. Play to them and the model does work you would otherwise do in an edit.
Use the long take as structure, not just length
Thirty seconds is only an advantage if the shot earns it. A long take should have a beginning, a turn, and an end, the way a real oner does.
- Write the arc, not the subject. "The camera holds on the empty platform, a train arrives, she steps off and walks past lens" gives the model somewhere to go for the full duration.
- Put the reveal late. Long takes are worth using when something changes partway through: light shifting, a door opening, a crowd clearing.
- Move the camera once, deliberately. One sustained push or pull across thirty seconds reads as craft; three moves in the same clip reads as indecision.
- Let a beat breathe. A held moment before the turn is what separates a scene from a demo reel, and short models never give you room for it.
Use references to build a series, not a single shot
Reference control pays off across shots, so the real win is planning a set of clips that belong together.
- Fix your cast before you generate. Lock the character, product, and location references first, then write each shot against that fixed set.
- Label every reference in the prompt. "The courier from the character reference, on the street from the location reference" removes the guesswork the model would otherwise fill in.
- Keep references clean over numerous. A few well-lit, uncluttered images beat a pile of busy ones, and a video reference of roughly 5 to 10 seconds is usually enough to establish a subject.
- Give separate angles, not one collage, when a subject has to be recognizable from more than one side.
Use in-pass audio to drive the cut
Because sound is generated with the picture, pacing belongs in the prompt rather than the timeline.
- Name the sound bed. "Low rain under it, no music" sets a mood the picture will match.
- Place the beat. Saying a door slams as the camera reaches the doorway gets picture and sound landing together.
- Ask for the silence too. A held quiet before a line or an impact is a directing choice, and it is easier to request than to cut in later.
Use prompt planning by writing in priority order
The model reasons through the brief before rendering, so the order you write in shapes what it protects.
- Lead with the shot, then the detail: subject and camera first, then wardrobe, light, and set dressing.
- State what must stay fixed. "The logo stays legible throughout" gives the planning step something to defend across the whole take.
- Give light a direction. "Low key from screen left" is something a plan can act on; "moody" is not.
- Keep one action per take. Three sequential beats in one prompt gets a compromise; three prompts get three clean shots.
For the full specification list, see the Wan 3.0 model page.
Wan 3.0 prompt guide
A strong video prompt reads like a short shot brief, not a caption, so the model has a subject, a motion, and a camera to work with rather than a still frame in words. Run through SPACE before you send.
| SPACE | Include | Example |
|---|---|---|
| Subject | Who or what is in frame, described concretely | A courier in a soaked yellow jacket |
| Performance | The motion: what the subject does, and how | He shoulders the door open and steps through |
| Ambience | Setting, time of day, and light | A narrow alley at night, neon spill on wet brick |
| Camera | Shot type plus one move | Mid shot, a slow push-in |
| Extra cues | Audio, pacing, and transitions | Rain bed under it, one unbroken take |
Weak vs strong prompts
Each row below turns a generic prompt into one that gives a specific Wan 3.0 strength something to work with.
| Strength in play | Weak | Strong |
|---|---|---|
| Long take | A violinist playing | She finishes the phrase, lowers the bow, and holds still as the last note decays and the camera drifts left |
| Camera control | A rainy street at night | Mid shot on a courier, one slow push-in down a neon-lit alley, rain falling through the key light |
| Reference continuity | Use these references | The courier from the character reference crosses the plaza shown in the location reference |
| In-pass audio | Add some sound | Low rain bed, no music, and the shutter slams as the camera reaches the doorway |
| Prompt planning | A product video with our logo | Studio-lit rotation of a brushed steel kettle, logo on the body legible throughout the turn |
Common mistakes
- Describing a still. A video model needs motion over time, not a photograph in words.
- Asking for thirty seconds with nothing happening in them. Length without a turn is just a slow clip.
- Writing "cinematic" and stopping. Name the shot type and one camera move instead.
- Cramming a sequence into one prompt. Keep one clear action per take and use references to carry continuity between them.
- Leaving references unlabeled. Say what each reference is for, or the model has to guess which one drives the scene.
- Treating audio as an afterthought. It generates with the picture, so unrequested sound is a missed choice rather than a neutral one.

