Bernini
by ByteDance
ByteDance's open‑source model for instruction‑based video edits, identity held.

Key features
Technical specifications
Planner + DiT
Qwen2.5-VL planner, 14B Wan2.2 renderer
Edit, Generate, R2V
Editing, generation, subject-to-video
480p / 16fps
Default render setting
Apache 2.0
Open weights, self-hostable
Use cases
Augment real footage
Add or remove props, fix a detail, or restyle an element in a clip without re-shooting. The consistency lock keeps the rest of the shot identical, so edits read as native.
Recurring characters and avatars
Keep the same face across episodes, ads, or an avatar series. Subject-to-video holds a person's identity from a few reference images as they move through new scenes.
Virtual try-on and product placement
Swap clothing onto a moving model from a reference, or drop a product or on-screen video into a shot, for fashion and ad work that needs the source clip kept intact.
Re-block an action
Change what someone is doing in a take, a stand becomes a crouch, without re-filming. Motion editing alters the action while identity, framing, and lighting stay fixed.
Prompt examples
Consistency edit
Add a snowman beside the dog on the snowy path, and keep the dog, the road, and the trees unchanged
Edit promptIdentity-locked subject
Place this person on a neon city rooftop at night, gently turning to camera, keeping their face and jacket
Edit promptReference swap
Replace the outer shirt with the one in the reference image, keep the pose, lighting, and motion exactly
Edit promptOverview
Bernini is ByteDance's open-source framework for AI video generation and editing, released under Apache 2.0 in June 2026. Its design splits the job in two: an MLLM-based planner (Qwen2.5-VL) decides the semantics of the result, then a DiT-based renderer built on Wan2.2 paints the pixels. Because the plan and the render are separate stages, Bernini is unusually good at the thing most video models struggle with, editing a real clip while leaving everything you didn't ask to change untouched. It spans text-to-image, image editing, text-to-video, instruction-based video editing, reference-guided edits like garment swaps and video insertion, and subject-to-video, where its standout result is preserving a person's identity as they move through a new scene.
Simple pricing
Get started for free today, with the option to upgrade or cancel anytime.
Basic
2400 monthly credits
1 user only
All models
Workflows
Standard
3625 monthly credits
1 user only
All models
Workflows
Pro
6350 shared monthly credits
1 user
All models
Workflows
Max
24650 shared monthly credits
1 user
All models
Workflows
Enterprise
For higher limits
Custom
pricing and billing terms

Free
For playing around
$0
forever free
FAQs
More about Bernini
Guides, prompts, and comparisons for Bernini.
Other models
Explore the rest of the Morphic model catalog.
Flux 3 Image
Black Forest Labs
Black Forest Labs' control-first image model. Place every element, edit one box at a time.
Ideogram 4.5
Ideogram
Ideogram's precision edit model. Stack edit on edit, with no drift or artifacts.
Eleven v4
ElevenLabs
ElevenLabs' most expressive voice model. Audio tags, 90+ languages, and a faster Turbo tier.
Kling 4.0 Flash
Kling
Kuaishou's speed-tuned Kling 4.0 for high-volume video. Fast 3 to 20 second clips at 720p.