Video generation

Vidu Q4

by Shengshu Technology

Shengshu's latest video model.
Up to 12 reference images and 3 voices held consistent across one clip, with native audio from 540p to 4K.

Vidu Q4

Key features

Technical specifications

Vidu Q

Shengshu's Vidu line, with Q4 the latest generation

2 modes

Image to video and reference to video, both live

3–16s

Any whole-second length in one pass (default 5s)

540p–4K

540p, 720p, 1080p, 2K, 4K (default 720p)

Up to 12

Characters, products, objects and locations held consistent

Up to 3

MP3 voice clips, 3 to 12 seconds each

Native

Synchronized dialogue and sound effects, generated with the clip

5 ratios

16:9, 9:16, 4:3, 3:4 and 1:1 on reference to video

Use cases

Character-driven storytelling

Tag your cast once and Vidu Q4 holds each face, outfit, and voice across the shot, so a character stays recognisable from the first frame to the last.

Branded content and ads

Drop in product shots and a spokesperson, keep them on-model across the clip, and let the scene speak with native audio — made for short ads and social spots.

Animating a still

Feed one image as the first frame and direct motion, camera moves, and a line of dialogue from the prompt, while the source subject and composition stay intact.

Multi-subject scenes

Combine characters, objects, and a location in a single prompt by referring to each image by its position, and the model composes them into one coherent scene.

Voiced dialogue

Supply up to three reference voice clips and point each line of dialogue at the voice that should say it, so speech stays consistent across the clip.

Multi-shot sequences

Block camera cuts within one generation, holding your referenced subjects across each cut rather than stitching separate clips together.

Prompt examples

Reference scene with a voice

[@reference_image_1] walks along the beach at sunset holding [@reference_image_2], then turns to the camera and speaks in the voice of [reference_audio_1]

Edit prompt

Animate a still with a line

The woman turns toward the camera and smiles, then the shot cuts to a close-up as she says: 'Let's go'

Edit prompt

Product with a spokesperson

[@reference_image_1] holds up [@reference_image_2] in a bright studio and says, in the voice of [reference_audio_1], 'Meet the new one'

Edit prompt

Multi-subject scene

[@reference_image_1] and [@reference_image_2] sit across [@reference_image_3], a small cafe table, talking as the camera slowly pushes in

Edit prompt

Multi-shot sequence

[@reference_image_1] in a red jacket: one shot weaving through a night market, then a hard cut to the same character arriving at a lit doorway at dusk

Edit prompt

Two voices in conversation

[@reference_image_1] and [@reference_image_2] talk across a table — the first speaks in the voice of [reference_audio_1], the second answers in [reference_audio_2]

Edit prompt

Overview

Vidu Q4 is the latest video generation model from Shengshu Technology. It builds a clip from references you supply — up to 12 images and up to 3 voice clips — or animates a single first frame, generating 3 to 16 seconds at up to 4K with optional native audio. It is the newest Vidu on Morphic.

What Vidu Q4 does differently

The headline is reference capacity. Where many video models take a single image, Vidu Q4 accepts up to 12 reference images in one prompt — characters, products, objects, locations — plus up to three voice clips, and keeps each subject and voice consistent. You point to each one by its position in the prompt, so one generation can assemble a cast, a set, and the right voices together.

Image-to-video and reference-to-video

Vidu Q4 runs in two modes. Image-to-video animates a single first frame, holding its subject and composition while the prompt directs motion, camera moves, and dialogue. Reference-to-video is the fuller mode: tag up to 12 images and up to 3 voice clips in the prompt, and the model composes them into one scene, keeping each character, product, and voice consistent. Audio is native in both and generated in the same pass as the video, and on reference-to-video you can toggle it off for a silent clip.

Vidu Q4 on Morphic

Vidu Q4 runs on Morphic, beside video models like Veo, Kling, Seedance, and the earlier Vidu releases. One set of references can go to two models on the same Canvas, which is the fastest way to find out whether Vidu Q4 holds your characters better than the alternative rather than taking a launch chart's word for it.

Simple pricing

Get started for free today, with the option to upgrade or cancel anytime.

Basic

$19/ month
billed as $0 per year

2400 monthly credits

1 user only

All models

Workflows

Standard

$24/ month
billed as $0 per year

3625 monthly credits

1 user only

All models

Workflows

Pro

$45/ month
billed as $0 per year

6350 shared monthly credits

1 user

+ up to 4 more at extra cost

All models

Workflows

Max

$170/ month
billed as $0 per year

24650 shared monthly credits

1 user

+ up to 9 more at extra cost

All models

Workflows

Enterprise

For higher limits

Custom

pricing and billing terms

High-volume credits
Custom seat limits
All models
Workflows
Pricing Gradient

Free

For playing around

$0

forever free

Up to 20 credits
1 user only
Limited models
Workflows

FAQs

What is Vidu Q4?
Vidu Q4 is the latest generation of Vidu, the video model from Shengshu Technology. It generates 3-to-16-second clips at up to 4K, from either a single first-frame image or up to 12 reference images, with optional native audio and up to three reference voice clips.
What's the difference between image-to-video and reference-to-video?
Image-to-video animates a single first-frame image, preserving its subject and composition while the prompt directs motion, cuts, and dialogue. Reference-to-video builds a scene from up to 12 reference images and up to 3 voice clips that you tag in the prompt, keeping each subject and voice consistent.
How many reference images and voice clips can Vidu Q4 use?
Up to 12 reference images (PNG, JPEG, JPG, or WebP) and up to 3 voice clips (MP3, 3 to 12 seconds each, up to 50 MB) on the reference-to-video mode. You tag each one in the prompt and the model keeps that subject or voice consistent across the clip.
Does Vidu Q4 generate audio?
Yes. It produces native, synchronized dialogue and sound effects in the same pass as the video, so sound is tied to what happens on screen rather than added afterward. On reference-to-video, audio is an optional toggle, so you can also render a silent clip.
How long are Vidu Q4 videos, and what resolution?
Any whole-second length from 3 to 16 seconds (default 5 seconds), at 540p, 720p, 1080p, 2K, or 4K (default 720p). Start low for fast iteration and step up to 4K for final delivery.
What aspect ratios does Vidu Q4 support?
Reference-to-video supports five: 16:9, 9:16, 4:3, 3:4, and 1:1, so one model covers a landscape frame and a vertical phone cut. Image-to-video follows the shape of the first-frame image you provide.
How do I keep a specific character or voice consistent?
Reference each one by its position in the prompt — the first image, the second image, and so on for images, and the first voice clip, the second, and so on for voices — and Vidu Q4 holds that subject or voice through the clip.
How is Vidu Q4 different from earlier Vidu models?
Vidu Q4 is the latest generation in the Vidu line from Shengshu Technology. The additions that stand out are reference capacity — up to 12 reference images and 3 voice clips held consistent in a single clip — alongside native synchronized audio and resolution up to 4K. On Morphic it runs beside the earlier Vidu releases, so you can send the same prompt to both and compare the result directly rather than relying on version numbers.
How do I use Vidu Q4 on Morphic?
Open Copilot, attach your reference images or a first frame, add any voice clips, describe the scene, and select Vidu Q4. Because Morphic runs it beside models like Veo, Kling, and Seedance, you can send the same references to two models on one Canvas and keep whichever take holds your characters best.