Vidu Q4
by Shengshu Technology
Shengshu's latest video model.
Up to 12 reference images and 3 voices held consistent across one clip, with native audio from 540p to 4K.

Key features
Technical specifications
Vidu Q
Shengshu's Vidu line, with Q4 the latest generation
2 modes
Image to video and reference to video, both live
3–16s
Any whole-second length in one pass (default 5s)
540p–4K
540p, 720p, 1080p, 2K, 4K (default 720p)
Up to 12
Characters, products, objects and locations held consistent
Up to 3
MP3 voice clips, 3 to 12 seconds each
Native
Synchronized dialogue and sound effects, generated with the clip
5 ratios
16:9, 9:16, 4:3, 3:4 and 1:1 on reference to video
Use cases
Character-driven storytelling
Tag your cast once and Vidu Q4 holds each face, outfit, and voice across the shot, so a character stays recognisable from the first frame to the last.
Branded content and ads
Drop in product shots and a spokesperson, keep them on-model across the clip, and let the scene speak with native audio — made for short ads and social spots.
Animating a still
Feed one image as the first frame and direct motion, camera moves, and a line of dialogue from the prompt, while the source subject and composition stay intact.
Multi-subject scenes
Combine characters, objects, and a location in a single prompt by referring to each image by its position, and the model composes them into one coherent scene.
Voiced dialogue
Supply up to three reference voice clips and point each line of dialogue at the voice that should say it, so speech stays consistent across the clip.
Multi-shot sequences
Block camera cuts within one generation, holding your referenced subjects across each cut rather than stitching separate clips together.
Prompt examples
Reference scene with a voice
[@reference_image_1] walks along the beach at sunset holding [@reference_image_2], then turns to the camera and speaks in the voice of [reference_audio_1]
Edit promptAnimate a still with a line
The woman turns toward the camera and smiles, then the shot cuts to a close-up as she says: 'Let's go'
Edit promptProduct with a spokesperson
[@reference_image_1] holds up [@reference_image_2] in a bright studio and says, in the voice of [reference_audio_1], 'Meet the new one'
Edit promptMulti-subject scene
[@reference_image_1] and [@reference_image_2] sit across [@reference_image_3], a small cafe table, talking as the camera slowly pushes in
Edit promptMulti-shot sequence
[@reference_image_1] in a red jacket: one shot weaving through a night market, then a hard cut to the same character arriving at a lit doorway at dusk
Edit promptTwo voices in conversation
[@reference_image_1] and [@reference_image_2] talk across a table — the first speaks in the voice of [reference_audio_1], the second answers in [reference_audio_2]
Edit promptOverview
Vidu Q4 is the latest video generation model from Shengshu Technology. It builds a clip from references you supply — up to 12 images and up to 3 voice clips — or animates a single first frame, generating 3 to 16 seconds at up to 4K with optional native audio. It is the newest Vidu on Morphic.
What Vidu Q4 does differently
The headline is reference capacity. Where many video models take a single image, Vidu Q4 accepts up to 12 reference images in one prompt — characters, products, objects, locations — plus up to three voice clips, and keeps each subject and voice consistent. You point to each one by its position in the prompt, so one generation can assemble a cast, a set, and the right voices together.
Image-to-video and reference-to-video
Vidu Q4 runs in two modes. Image-to-video animates a single first frame, holding its subject and composition while the prompt directs motion, camera moves, and dialogue. Reference-to-video is the fuller mode: tag up to 12 images and up to 3 voice clips in the prompt, and the model composes them into one scene, keeping each character, product, and voice consistent. Audio is native in both and generated in the same pass as the video, and on reference-to-video you can toggle it off for a silent clip.
Vidu Q4 on Morphic
Vidu Q4 runs on Morphic, beside video models like Veo, Kling, Seedance, and the earlier Vidu releases. One set of references can go to two models on the same Canvas, which is the fastest way to find out whether Vidu Q4 holds your characters better than the alternative rather than taking a launch chart's word for it.
Simple pricing
Get started for free today, with the option to upgrade or cancel anytime.
Basic
2400 monthly credits
1 user only
All models
Workflows
Standard
3625 monthly credits
1 user only
All models
Workflows
Pro
6350 shared monthly credits
1 user
All models
Workflows
Max
24650 shared monthly credits
1 user
All models
Workflows
Enterprise
For higher limits
Custom
pricing and billing terms

Free
For playing around
$0
forever free
FAQs
Nano Banana 2.1
Google DeepMind
Google DeepMind's new Flash image model. Better design, cleaner edits, steadier characters.
Flux 3 Image
Black Forest Labs
Black Forest Labs' control-first image model. Place every element, edit one box at a time.
Ideogram 4.5
Ideogram
Ideogram's precision edit model. Stack edit on edit, with no drift or artifacts.
Eleven v4
ElevenLabs
ElevenLabs' most expressive voice model. Audio tags, 90+ languages, and a faster Turbo tier.