Flux 3
by Black Forest Labs
Black Forest Labs' world model for image, video, and audio, with sound that matches the action.

Key features
Technical specifications
5–20s
Whole seconds in a single pass, or automatic to suit the prompt
720p / 1080p
HD by default, Full HD through upscaling
T2V·I2V·V2V
Text, image and keyframes, or continuation of an existing clip
Native
Synchronized audio generated with every clip, on by default
Use cases
Video with sound in one pass
A prompt returns a clip and its audio together, so an effect or a line is already synced to the action instead of added later.
Multi-shot sequences
Block several angles in one generation, or extend a clip with continuation, holding one character across every cut.
Product and brand film
Materials, weight, and reflections stay honest as the camera moves, so a product reads as filmed rather than simulated.
Localized video
Multilingual dialogue and accurate on-screen text let one scene ship in several languages without a separate title or voice pass.
Choreographed transitions
Pin the poses and compositions a shot has to move through, and the model builds the motion between them as one continuous take.
Explainers with real physics
Gravity, collisions, and motion follow real rules, so science and how-it-works footage stays accurate while it stays clean.
Prompt examples

Physics-true clip
A ceramic mug slips off a marble counter and shatters on the floor, shards scattering, sharp impact sound in sync
Edit prompt
Multi-shot sequence
Same courier in a red jacket: one shot weaving through a market, the next arriving at a lit doorway at dusk
Edit prompt
Localized title card
Coffee brand hero shot, steam rising off a fresh pour, on-screen title reading 'Freshly Roasted' in crisp bold type
Edit prompt
Camcorder to cinematic
Handheld camcorder footage of a birthday party, then the same scene as a warm cinematic wide with room tone
Edit prompt
Clip continuation
She turns from the window and walks out into the rain-lit Tokyo street, the camera following without a cut
Edit prompt
Studio product
Avant-garde sneaker rotating on a titanium plinth, hard key light, accurate metal reflections, faint mechanical hum
Edit promptOverview
Flux 3 is Black Forest Labs' multimodal foundation model, announced on July 23, 2026, with video generation made generally available on August 4, 2026. It jointly learns from images, video, and audio within a single unified architecture, so one model can render a moving shot and the sound that belongs to it, holding a coherent sense of how the world looks, moves, and sounds.
What Flux 3 does differently
Most generators treat each modality as a separate problem. Flux 3 treats image, video, and audio as evidence about one underlying reality, trained together on Black Forest Labs' Self-Flow approach, with over 95% of total training compute spent on video prediction. The payoff shows up as physical consistency: the sound matches the impact, the motion obeys the mass, and a subject keeps its identity from one shot to the next. That same world understanding is what lets the backbone extend to action prediction for robotics, tested on real production tasks at Audi through the FLUX-mimic partnership.
What has shipped, and what has not
Flux 3 Video is the generally available piece: text-to-video, image-to-video with keyframes, and continuation of an existing clip, all at up to 20 seconds with synchronized audio. Black Forest Labs has also announced Flux 3 Image for image generation and editing, Flux 3 Action for robotics, and an open-weight Flux 3 Dev, plus video editing and combined image, video, and audio references. Those are on the roadmap rather than in your hands today.
Flux 3 on Morphic
Flux 3 runs on Morphic, beside video models like Veo, Kling, Seedance, and Vidu, and the rest of the Flux family for images. One prompt can go to two models on the same Canvas, which is the fastest way to find out whether Flux 3 is the right call for a given shot rather than taking a launch chart's word for it.
Simple pricing
Get started for free today, with the option to upgrade or cancel anytime.
Basic
1100 monthly credits
1 user only
All models
Workflows
Standard
3625 monthly credits
1 user only
All models
Workflows
Pro
6350 shared monthly credits
1 user
All models
Workflows
Pro Max
24650 shared monthly credits
1 user
All models
Workflows
Enterprise
For higher limits
Custom
pricing and billing terms

Free
For playing around
$0
forever free
FAQs
ChatGPT Images 2.5
OpenAI
OpenAI's image model, released September 2026. Sharper detail, precise edits, faster.
Lyria 3.5
Google DeepMind
Google's best-sounding music model. Full songs with structure, vocals, and lyrics.
MiniMax H3 Max Turbo
fal.ai
The fast tier of fal's post-trained H3. Twice the speed, nearly the same look.
Inworld TTS 2
Inworld
Ninety-five voices, 100+ languages. Speech that starts in under a second.