Multi-modal

Flux 3

by Black Forest Labs

Black Forest Labs' world model for image, video, and audio, with sound that matches the action.

Flux 3

Key features

Technical specifications

Image·Video·Audio

Learned jointly in one model; extends to action prediction

Up to 20s

Single-generation video length; chain clips for longer sequences

Native

Synchronized audio generated with every video clip

Self-Flow

Unified multimodal architecture from Black Forest Labs

Use cases

Video with sound in one pass

A prompt returns a clip and its audio together, so an effect or a line is already synced to the action instead of added later.

Multi-shot sequences

Chain clips into a longer scene and hold one character across every cut with a visual reference, from a short to a long piece.

Product and brand film

Materials, weight, and reflections stay honest as the camera moves, so a product reads as filmed rather than simulated.

Localized video

Multilingual dialogue and accurate on-screen text let one scene ship in several languages without a separate title or voice pass.

Image generation and editing

Synthesize across styles and aspect ratios, render crisp multilingual text, then edit one region in place without a full re-roll.

Explainers with real physics

Gravity, collisions, and motion follow real rules, so science and how-it-works footage stays accurate while it stays clean.

Prompt examples

Physics-true clip

Physics-true clip

A ceramic mug slips off a marble counter and shatters on the floor, shards scattering, sharp impact sound in sync

Edit prompt
Multi-shot sequence

Multi-shot sequence

Same courier in a red jacket: one shot weaving through a market, the next arriving at a lit doorway at dusk

Edit prompt
Localized title card

Localized title card

Coffee brand hero shot, steam rising off a fresh pour, on-screen title reading 'Freshly Roasted' in crisp bold type

Edit prompt
Camcorder to cinematic

Camcorder to cinematic

Handheld camcorder footage of a birthday party, then the same scene as a warm cinematic wide with room tone

Edit prompt
In-context edit

In-context edit

Take the subject from the reference clip and set them walking a rain-lit Tokyo street at night, keep the rest

Edit prompt
Studio product

Studio product

Avant-garde sneaker rotating on a titanium plinth, hard key light, accurate metal reflections, faint mechanical hum

Edit prompt

Simple pricing

Get started for free today, with the option to upgrade or cancel anytime.

Basic

$9/ month
billed as $0 per year

900 monthly credits

1 user only

All models

Workflows

Standard

$24/ month
billed as $0 per year

3200 monthly credits

1 user only

All models

Workflows

Pro

$45/ month
billed as $0 per year

6200 shared monthly credits

1 user

+ up to 4 more at extra cost

All models

Workflows

Pro Max

$170/ month
billed as $0 per year

24000 shared monthly credits

1 user

+ up to 9 more at extra cost

All models

Workflows

Enterprise

For higher limits

Custom

pricing and billing terms

High-volume credits
Custom seat limits
All models
Workflows
Pricing Gradient

Free

For playing around

$0

forever free

Up to 20 credits
1 user only
Limited models
Workflows

FAQs

What is Flux 3?
Flux 3 is Black Forest Labs' multimodal foundation model, announced on July 23, 2026. It jointly learns from images, video, and audio in one unified architecture, so a single model can generate video with native audio, synthesize and edit images, and hold a coherent understanding of how the real world looks, moves, and sounds.
Is Flux 3 an image model or a video model?
Both. Flux 3 is one multimodal model rather than separate tools. It generates and edits images, and it generates video with synchronized native audio, all from the same world model. Black Forest Labs also extends the backbone to action prediction for robotics through its FLUX-mimic partnership, tested on real production tasks at Audi.
How long are Flux 3 videos?
Flux 3 generates up to 20 seconds of video in a single pass. For longer pieces, it chains clips into multi-shot sequences and reuses a visual reference to keep a character consistent across scenes that can run for several minutes. Each clip returns with synchronized native audio already matched to the action.
Does Flux 3 generate audio?
Yes. Every Flux 3 video is generated with native audio in the same pass, so effects, ambience, and dialogue are already synced to what happens on screen. Sound is tied to physical events, so an impact matches the hit, and Flux 3 supports multilingual dialogue, letting a scene speak in more than one language.
Can Flux 3 edit images and video?
Yes. Flux 3 does in-context editing for both. You can change one element of an image and leave the rest of the frame intact, or carry a subject from a source video into a new scene. It builds on the instruction-editing lineage of Flux Kontext, which already powers image editing on Morphic.
How is Flux 3 different from Flux 2?
Flux 2 Pro is an image model. Flux 3 is a multimodal foundation model that adds native-audio video, motion that follows real physics, and one shared representation across image, video, and audio. Text rendering and prompt handling also improve significantly, so Flux 3 is a world model where Flux 2 focused on still images.
Is Flux 3 open source?
Black Forest Labs has said it will release Flux 3 Dev, an open-weight multimodal backbone covering image, video, audio, and action prediction, alongside faster versions. The open-weight release is planned to follow the early-access rollout of the API and private-weight models, so open weights are coming rather than available at launch.
How do I use Flux 3 on Morphic?
The Flux family already powers image work on Morphic through Flux 2 Pro and Kontext editing, alongside models like Nano Banana Pro and Seedream 5. Open Morphic, describe what you want, and generate on the Canvas, comparing models side by side so you keep the render that lands best for your shot.