Multi-modal

Flux 3

by Black Forest Labs

Black Forest Labs' world model for image, video, and audio, with sound that matches the action.

Flux 3

Key features

Technical specifications

5–20s

Whole seconds in a single pass, or automatic to suit the prompt

720p / 1080p

HD by default, Full HD through upscaling

T2V·I2V·V2V

Text, image and keyframes, or continuation of an existing clip

Native

Synchronized audio generated with every clip, on by default

Use cases

Video with sound in one pass

A prompt returns a clip and its audio together, so an effect or a line is already synced to the action instead of added later.

Multi-shot sequences

Block several angles in one generation, or extend a clip with continuation, holding one character across every cut.

Product and brand film

Materials, weight, and reflections stay honest as the camera moves, so a product reads as filmed rather than simulated.

Localized video

Multilingual dialogue and accurate on-screen text let one scene ship in several languages without a separate title or voice pass.

Choreographed transitions

Pin the poses and compositions a shot has to move through, and the model builds the motion between them as one continuous take.

Explainers with real physics

Gravity, collisions, and motion follow real rules, so science and how-it-works footage stays accurate while it stays clean.

Prompt examples

Physics-true clip

Physics-true clip

A ceramic mug slips off a marble counter and shatters on the floor, shards scattering, sharp impact sound in sync

Edit prompt
Multi-shot sequence

Multi-shot sequence

Same courier in a red jacket: one shot weaving through a market, the next arriving at a lit doorway at dusk

Edit prompt
Localized title card

Localized title card

Coffee brand hero shot, steam rising off a fresh pour, on-screen title reading 'Freshly Roasted' in crisp bold type

Edit prompt
Camcorder to cinematic

Camcorder to cinematic

Handheld camcorder footage of a birthday party, then the same scene as a warm cinematic wide with room tone

Edit prompt
Clip continuation

Clip continuation

She turns from the window and walks out into the rain-lit Tokyo street, the camera following without a cut

Edit prompt
Studio product

Studio product

Avant-garde sneaker rotating on a titanium plinth, hard key light, accurate metal reflections, faint mechanical hum

Edit prompt

Overview

Flux 3 is Black Forest Labs' multimodal foundation model, announced on July 23, 2026, with video generation made generally available on August 4, 2026. It jointly learns from images, video, and audio within a single unified architecture, so one model can render a moving shot and the sound that belongs to it, holding a coherent sense of how the world looks, moves, and sounds.

What Flux 3 does differently

Most generators treat each modality as a separate problem. Flux 3 treats image, video, and audio as evidence about one underlying reality, trained together on Black Forest Labs' Self-Flow approach, with over 95% of total training compute spent on video prediction. The payoff shows up as physical consistency: the sound matches the impact, the motion obeys the mass, and a subject keeps its identity from one shot to the next. That same world understanding is what lets the backbone extend to action prediction for robotics, tested on real production tasks at Audi through the FLUX-mimic partnership.

What has shipped, and what has not

Flux 3 Video is the generally available piece: text-to-video, image-to-video with keyframes, and continuation of an existing clip, all at up to 20 seconds with synchronized audio. Black Forest Labs has also announced Flux 3 Image for image generation and editing, Flux 3 Action for robotics, and an open-weight Flux 3 Dev, plus video editing and combined image, video, and audio references. Those are on the roadmap rather than in your hands today.

Flux 3 on Morphic

Flux 3 runs on Morphic, beside video models like Veo, Kling, Seedance, and Vidu, and the rest of the Flux family for images. One prompt can go to two models on the same Canvas, which is the fastest way to find out whether Flux 3 is the right call for a given shot rather than taking a launch chart's word for it.

Simple pricing

Get started for free today, with the option to upgrade or cancel anytime.

Basic

$9/ month
billed as $0 per year

1100 monthly credits

1 user only

All models

Workflows

Standard

$24/ month
billed as $0 per year

3625 monthly credits

1 user only

All models

Workflows

Pro

$45/ month
billed as $0 per year

6350 shared monthly credits

1 user

+ up to 4 more at extra cost

All models

Workflows

Pro Max

$170/ month
billed as $0 per year

24650 shared monthly credits

1 user

+ up to 9 more at extra cost

All models

Workflows

Enterprise

For higher limits

Custom

pricing and billing terms

High-volume credits
Custom seat limits
All models
Workflows
Pricing Gradient

Free

For playing around

$0

forever free

Up to 20 credits
1 user only
Limited models
Workflows

FAQs

What is Flux 3?
Flux 3 is Black Forest Labs' multimodal foundation model, announced on July 23, 2026. It jointly learns from images, video, and audio in one unified architecture, so a single model can generate video with synchronized audio and hold a coherent understanding of how the real world looks, moves, and sounds. Video generation became generally available on August 4, 2026.
Is Flux 3 an image model or a video model?
It is one multimodal model, and video is the part that has shipped. Black Forest Labs spent over 95% of its total training compute on video prediction, and Flux 3 Video is generally available now. Flux 3 Image, Flux 3 Action for robotics, and an open-weight Flux 3 Dev are announced and still rolling out.
How long are Flux 3 videos?
Flux 3 generates 5 to 20 seconds in a single pass, in whole seconds, or it can choose a length to suit the prompt. For longer pieces, continuation extends a clip from its final frames, and a single generation can block several shots and camera angles with hard cuts between them.
Does Flux 3 generate audio?
Yes, and it is on by default. Every clip is generated with synchronized audio in the same pass, so effects, ambience, and dialogue are already tied to what happens on screen. Sound follows physical events, so an impact matches the hit. You can also turn audio off for a silent clip.
What video modes does Flux 3 support?
Three. Text-to-video works from a prompt alone. Image-to-video takes stills, either as the opening frame or as up to ten keyframes pinned to timestamps. Video continuation extends an existing clip from its final frames. Video editing and combined image, video, and audio references are announced as coming next.
What resolution and aspect ratios does Flux 3 support?
Flux 3 renders at HD 720p by default, with Full HD 1080p available through upscaling. Seven aspect ratios are supported: 21:9, 2:1, 16:9, 4:3, 1:1, 3:4, and 9:16, so one model covers a cinemascope frame and a vertical phone cut.
How is Flux 3 different from Flux 2?
Flux 2 Pro is an image model. Flux 3 is a multimodal foundation model that adds video with synchronized audio, motion that follows real physics, and one shared representation across image, video, and audio. It is Black Forest Labs' first video model, where the Flux 2 line focused on still images.
How good is Flux 3 compared to other video models?
In Black Forest Labs' own human-rater evaluation at release, Flux 3 led text-to-video against every model it was tested against, and tied Seedance 2.0 while beating the rest in image-to-video. Those are the developer's internal results, so treat them as a strong signal rather than an independent benchmark.
Is Flux 3 open source?
Not yet. Black Forest Labs has said it will release Flux 3 Dev, an open-weight multimodal backbone covering image, video, audio, and action prediction. The open-weight release is planned to follow the API and private-weight rollout, so open weights are announced rather than available today.
How do I use Flux 3 on Morphic?
Flux 3 runs on Morphic today. Open the Canvas, describe the shot with one line about how it should sound, and generate. Because Morphic runs Flux 3 beside models like Veo, Kling, and Seedance, you can send the same prompt to two of them and keep whichever take lands better.