Wan 3.0 guide: features, prompts, and worked examples

Wan 3.0 guide: features, prompts, and worked examples

A Wan 3.0 guide to its real features, native 30-second takes, document-to-video, and reference control, with prompt examples and the habits that hold quality.

Wan 3.0 features and capabilities

Wan 3.0 is Alibaba Tongyi Lab's latest video model, in public beta since August 2026. It generates native 1080p clips up to 30 seconds in a single run, reads documents and webpages as well as prompts, and generates audio in the same pass as the picture.

CapabilityWhat it doesBest for
Long single takesHolds one continuous shot for up to 30 secondsAd spots, short scenes, long reveals
Video from documentsReads a doc, sheet, deck, pdf, or webpage and builds video from itReports, explainers, data-driven clips
Omni-ReferenceHolds a character, product, or set across shots, from up to 20 referencesSeries, multi-shot sequences, brand work
Audio in the same passReturns a scored clip rather than a silent oneSound-on scenes, dialogue, music beats
Precision editingReselect a time interval and regenerate only that partFixes and revisions without a full reroll

Long single takes

Length is the headline. Up to 30 seconds in one continuous shot is about double what the line did at Wan 2.7, and it changes the unit of work: a full ad spot or a short scene fits inside one generation instead of being cut together from several. You write it as one evolving motion, describing how the subject and camera move across the whole take rather than a single frozen frame. Intelligent duration control matches the length to the brief, so a short beat is not padded out with dead seconds.

Video from documents

This is what most sets Wan 3.0 apart. Its Omni-Reference input reads structured files, doc, xls, ppt, pdf, txt, key, pages, numbers, and md, plus any webpage URL, and builds a video sequence from the content inside. A quarterly deck or a product spec sheet becomes a finished clip rather than a storyboard. Give it clean, well-organized source, then say in the prompt which figures, lines, or sections should carry the video.

Reference-led continuity

Omni-Reference is what keeps a sequence recognizable. Supply up to 20 assets, images, clips, audio, even documents, of a character, product, or set, and the model carries that identity through the shot and from cut to cut. Naming what each reference is for matters as much as supplying it, so the model knows which one drives the subject and which fixes the location.

Audio in the same pass

Sound generates alongside the picture rather than being scored afterwards, so a scene can come back with room tone, effects, or dialogue already in place. That makes pacing part of the prompt: say where a beat lands and the picture can be cut to it.

Precision editing

Editing lives in the same model. Wan 3.0 supports instruction- and reference-based editing: select a time interval and regenerate only that part, leaving the rest untouched, and adjust the visuals, the action, or a character's lines. It turns revision into a controllable pass rather than rerolling a whole clip and hoping for a better draw.

On open weights

Wan 3.0 is a closed, API-only model, with no published weights on Hugging Face, GitHub, or ModelScope. Alibaba's open Wan line stops at Wan 2.2 under Apache 2.0, and those weights are still the ones you can download and self-host. If self-hosting is a requirement for your pipeline, that is the detail to watch, and Wan 2.2 remains the newest downloadable Wan.

How to get the best out of Wan 3.0

Each of Wan 3.0's strengths asks for something specific from the prompt. Play to them and the model does work you would otherwise do in an edit.

Use the long take as structure, not just length

Thirty seconds is only an advantage if the shot earns it. A long take should have a beginning, a turn, and an end, the way a real oner does.

  • Write the arc, not the subject. "The camera holds on the empty platform, a train arrives, she steps off and walks past lens" gives the model somewhere to go for the full duration.
  • Put the reveal late. Long takes are worth using when something changes partway through: light shifting, a door opening, a crowd clearing.
  • Move the camera once, deliberately. One sustained push or pull across thirty seconds reads as craft; three moves in the same clip reads as indecision.
  • Let a beat breathe. A held moment before the turn is what separates a scene from a demo reel, and short models never give you room for it.

Use references to build a series, not a single shot

Reference control pays off across shots, so the real win is planning a set of clips that belong together.

  • Fix your cast before you generate. Lock the character, product, and location references first, then write each shot against that fixed set.
  • Label every reference in the prompt. "The courier from the character reference, on the street from the location reference" removes the guesswork the model would otherwise fill in.
  • Keep references clean over numerous. A few well-lit, uncluttered images beat a pile of busy ones, and a video reference of roughly 5 to 10 seconds is usually enough to establish a subject.
  • Give separate angles, not one collage, when a subject has to be recognizable from more than one side.

Use in-pass audio to drive the cut

Because sound is generated with the picture, pacing belongs in the prompt rather than the timeline.

  • Name the sound bed. "Low rain under it, no music" sets a mood the picture will match.
  • Place the beat. Saying a door slams as the camera reaches the doorway gets picture and sound landing together.
  • Ask for the silence too. A held quiet before a line or an impact is a directing choice, and it is easier to request than to cut in later.

Write in priority order

A structured prompt gives the model a clear hierarchy to hold onto, so the order you write in shapes what it protects across a 30-second take.

  • Lead with the shot, then the detail: subject and camera first, then wardrobe, light, and set dressing.
  • State what must stay fixed. "The logo stays legible throughout" gives the planning step something to defend across the whole take.
  • Give light a direction. "Low key from screen left" is something a plan can act on; "moody" is not.
  • Keep one action per take. Three sequential beats in one prompt gets a compromise; three prompts get three clean shots.

For the full specification list, see the Wan 3.0 model page.

Wan 3.0 prompt guide

A strong video prompt reads like a short shot brief, not a caption, so the model has a subject, a motion, and a camera to work with rather than a still frame in words. Run through SPACE before you send.

SPACEIncludeExample
SubjectWho or what is in frame, described concretelyA courier in a soaked yellow jacket
PerformanceThe motion: what the subject does, and howHe shoulders the door open and steps through
AmbienceSetting, time of day, and lightA narrow alley at night, neon spill on wet brick
CameraShot type plus one moveMid shot, a slow push-in
Extra cuesAudio, pacing, and transitionsRain bed under it, one unbroken take

Weak vs strong prompts

Each row below turns a generic prompt into one that gives a specific Wan 3.0 strength something to work with.

Strength in playWeakStrong
Long takeA violinist playingShe finishes the phrase, lowers the bow, and holds still as the last note decays and the camera drifts left
Camera controlA rainy street at nightMid shot on a courier, one slow push-in down a neon-lit alley, rain falling through the key light
Reference continuityUse these referencesThe courier from the character reference crosses the plaza shown in the location reference
In-pass audioAdd some soundLow rain bed, no music, and the shutter slams as the camera reaches the doorway
In-frame textA product video with our logoStudio-lit rotation of a brushed steel kettle, logo on the body legible throughout the turn

Common mistakes

  • Describing a still. A video model needs motion over time, not a photograph in words.
  • Asking for thirty seconds with nothing happening in them. Length without a turn is just a slow clip.
  • Writing "cinematic" and stopping. Name the shot type and one camera move instead.
  • Cramming a sequence into one prompt. Keep one clear action per take and use references to carry continuity between them.
  • Leaving references unlabeled. Say what each reference is for, or the model has to guess which one drives the scene.
  • Treating audio as an afterthought. It generates with the picture, so unrequested sound is a missed choice rather than a neutral one.

Simple pricing

Get started for free today, with the option to upgrade or cancel anytime.

Basic

$9/ month
billed as $0 per year

900 monthly credits

1 user only

All models

Workflows

Standard

$24/ month
billed as $0 per year

3200 monthly credits

1 user only

All models

Workflows

Pro

$45/ month
billed as $0 per year

6200 shared monthly credits

1 user

+ up to 4 more at extra cost

All models

Workflows

Pro Max

$170/ month
billed as $0 per year

24000 shared monthly credits

1 user

+ up to 9 more at extra cost

All models

Workflows

Enterprise

For higher limits

Custom

pricing and billing terms

High-volume credits
Custom seat limits
All models
Workflows
Pricing Gradient

Free

For playing around

$0

forever free

Up to 20 credits
1 user only
Limited models
Workflows

FAQs

Is Wan 3.0 available to use yet?
Yes. Wan 3.0 has been in public beta since 6 August 2026 on Alibaba Cloud Model Studio and Qwen Cloud, as the model wan3.0-video. It is not in the Morphic catalog, so the prompting guidance below is written for Wan 3.0 and applies to any current video model, including the ones Morphic runs, in the meantime.
What does Wan 3.0 do?
Two things stand out: native 30-second clips in a single run, double Wan 2.7's 15 seconds, and Omni-Reference, which reads documents, spreadsheets, slides, pdfs, and webpages and builds video from them. Audio is generated in the same pass, and it outputs at 480p, 720p, or 1080p. There is no 4K tier, despite spec tables online that claim one.
Can Wan 3.0 turn a document into a video?
Yes, and it is the model's signature feature. Wan 3.0 accepts doc, xls, ppt, pdf, txt, key, pages, numbers, and md files plus webpage URLs, reads the content, and generates a video sequence from it. Supply clean, well-structured source and name in the prompt what you want emphasised, so the model knows which parts of the document carry the clip.
How do I write a good Wan prompt?
Treat the prompt as a short shot brief rather than a caption. Name the subject, the motion, the setting and light, and the camera move. Describe how the shot evolves across its length instead of a single frozen frame, and name a specific move like a slow push-in rather than writing "cinematic".
How do references work in Wan 3.0?
Omni-Reference lets you supply up to 20 assets, images, video, audio, even documents, and the model carries a character, product, or set through the clip and across shots. Say in the prompt what each reference is for, so the model knows which one drives the subject and which sets the location. Reusing the same references between cuts is what keeps a sequence consistent.
Will Wan 3.0 have open weights?
No. Wan 3.0 is a closed, API-only model with no weights published on Hugging Face, GitHub, or ModelScope. Alibaba's open Wan line stops at Wan 2.2 under Apache 2.0, so if you need a self-hosted Wan, that remains the newest one. Any site advertising Wan 3.0 weights or a free download is not offering the real model.