Flux 3: complete guide, features, prompts, and native-audio video

Flux 3: complete guide, features, prompts, and native-audio video

The complete Flux 3 guide: native-audio video up to 20 seconds, the four prompt formats, the CASTLE schema, keyframes, clip continuation, and worked examples.

Flux 3 features and capabilities

Flux 3 is Black Forest Labs' multimodal foundation model. Rather than a separate model per output, it learns image, video, and audio together, so a single prompt returns a moving, sounding shot that behaves the way the real world does.

Black Forest Labs announced Flux 3 on July 23, 2026 and made video generation generally available on August 4, 2026. Clips run 5 to 20 seconds at HD 720p or Full HD 1080p, in any of seven aspect ratios from 21:9 down to 9:16, with audio generated alongside the frames by default. Everything below follows Black Forest Labs' published prompting guide.

FeatureWhat it doesBest for
Native-audio videoGenerates the clip and its synchronized sound in one passSound-on scenes, dialogue, effects
Real-world physicsMotion obeys mass and momentum; sound matches impactProduct film, action, explainers
KeyframesInterpolates one shot through up to 10 pinned stillsControlled transitions, storyboards
Video continuationExtends an existing clip from its final framesLonger scenes, salvaging a good take
Multilingual dialogueSpeaks 13+ languages on camera with lip-syncLocalized video, presenters, ads
Multi-shot generationBlocks several angles inside a single generationShort films, ads, reveals
Draft modeFast preview, then a full render that matches itExploring directions before committing

Native-audio video

The headline is that picture and sound arrive together. Flux 3 generates each clip with synchronized native audio in the same pass, so an impact, footstep, or spoken line is already tied to the action instead of layered on after. Name the sound you want in the prompt, and set a language when the scene has dialogue.

Real-world physics

Flux 3 learns motion, mass, and sound as one system, so weight, momentum, and collisions follow real rules and materials keep their look as the camera moves. Describe how the motion evolves across the shot, and the physics reads as filmed rather than drifting.

Keyframes and continuation

Two ways to take control of the motion. Pin stills as keyframes and Flux 3 interpolates one continuous shot through them: a single image sets the opening frame, two fix a start and an end, and up to ten pinned to timestamps turn the clip into a storyboard the model has to hit in order. Or hand it footage you already have, and continuation picks up from the final frames, carrying momentum, camera behavior, and the audio bed across the seam without a cut.

Multilingual dialogue and typography

Quote the line, name the language, and describe the delivery, and Flux 3 speaks it on camera with matching accent and lip-sync. Supported languages include English and its dialects, Chinese, Spanish, French, German, Japanese, Portuguese, Russian, Italian, Indonesian, Turkish, Hindi, and Punjabi. Typography renders as part of the scene and holds up while the camera moves, so titles and signage survive the motion.

Flux 3 use cases

Native-audio scenes

A clip returns with its sound already synced, so an effect or a spoken line lands with the action instead of a silent take you fix later.

Product and brand film

Reflections, weight, and materials stay honest as the camera moves, so a pour, a rotation, or a drop reads as filmed rather than simulated.

Multi-shot sequences

Block several angles inside one generation, or extend a clip with continuation, holding one character across every cut.

The same red-haired woman held consistent across three scenes

Localized video

Multilingual dialogue and accurate on-screen text let one scene ship in several languages, no separate title or voice pass needed.

A broadcast lower-third reading the same phrase in English and Korean

Draft, then commit

Preview a direction fast and cheap, judge it, then render the one you approved at full quality with the same subjects, composition, and motion.

A fast draft preview beside the full-quality render of the same shot

Explainers with real physics

Gravity, collisions, and motion follow real rules, so how-it-works footage stays accurate while it stays clean.

How to prompt Flux 3

Direct a scene, do not describe a pile of objects. Black Forest Labs' guidance is to name what happens, how subjects move, how the camera behaves, and what the shot sounds like, using concrete nouns and verbs a camera could actually see. Vague adjectives leave the result to chance.

Pick a prompt format

Flux 3 accepts four, and length is not the goal. Start short to explore, then lengthen only where you need control.

FormatUse it whenExample
Short phraseExploring fast, one clear subject, happy accidentsA red fox leaping through fresh snow, telephoto
Natural-language one-linerThe everyday default for a single shotA low tracking shot of a fox sprinting through wet pine undergrowth at dawn
Labeled fieldsTuning one lever at a time without touching the restCamera shot: wide, low angle. Subject: a rider crosses a river
TimestepThe action has to land on a mark0.0–1.5s locked wide of a harbor. 1.5–3.0s a slow push-in begins

The one-liner has a reliable shape worth memorizing: [camera] shot of [subject] [action] in [environment], then the supporting visual and motion detail. Keep timestep beats achievable, two or three for a five-second clip, and mark a hard cut wherever the angle changes.

Build a CASTLE for multi-shot work

When a look and a story have to hold across several shots, Black Forest Labs documents a six-part schema. The six parts spell CASTLE, which is our shorthand for their schema rather than an official name, and it is the fastest way to remember what a long prompt is missing.

ElementWhat goes in it
CCore summaryOne line covering the whole sequence: who, where, and the arc
AAudioPer shot: the soundscape, rendered in sync with the frames
SSubjectHeld word-for-word identical across shots, so identity stays put
TTimelinePer shot, timecoded: the camera move and the subject's action
LLookThe global finish: realism level, palette anchors, and grain
EEnvironmentPer shot: setting, light quality, and depth of field

Fill them in whichever order you think, then assemble the prompt in Black Forest Labs' documented sequence: core summary, environment, subject, timeline, audio, look. Keeping the subject description byte-identical between shots is what stops a character drifting across a cut. It is the cheapest continuity trick in the guide.

Name the sound you want

Audio is generated with the picture, so a scene that implies sound gives the model more to work with than an abstract one. Footsteps, impacts, rain, engines, and crowds all read clearly. There are four layers, and you rarely need all four.

LayerWhat to describeExample
SpeechWho speaks, the exact words in quotes, and howThe mechanic says, "Try it now." Quiet, matter-of-fact
AmbienceThe sound of the placeRain against the windows, low diner chatter
EffectsSounds tied to visible actionsA ceramic mug clicks against the saucer
MusicStyle, pace, and where it sits in the mixA sparse piano cue under the scene, low in the mix

Two habits do most of the work. Name a sound source rather than asking for silence, because "quiet room tone" can collapse into dead air while "rain against the window" gives the model something to render. And for a spoken line, say whether the speaker is on camera or it is a voiceover, otherwise a quoted line can be treated as text to put in the frame.

Direct the voice, not the vibe

"Professional, warm, and engaging" produces the same polished announcer read every time. Concrete anchors work better: age and accent when they matter, register, how it was recorded, the delivery, and a guardrail such as "no announcer delivery." Then read the line aloud before you generate it. Voice direction cannot rescue stiff copy.

Common mistakes

  • Describing a still. A video model needs motion over time, not a photograph in words.
  • Forgetting the sound. Audio generates by default, so name the effect, ambience, or line you want.
  • Stacking camera terms. "Low tracking shot" is clear; "low aerial handheld orbit push-in" is not.
  • Writing negatives. Flux models do not support negative prompts, so state what you do want.
  • Overfilling a short clip. Several speakers, a long script, and multiple beats in five seconds will clip the last word. Shorten the line or raise the duration.

For the full capability list and specifications, see the Flux 3 model page.

FAQs

How do I write a good Flux 3 prompt?
Direct a scene rather than describing objects. Name the subject, what it does, how the camera behaves, the setting and light, and one audio cue, using concrete nouns and verbs a camera could see. Because Flux 3 generates sound with the picture, always say what the shot should sound like.
What are the four Flux 3 prompt formats?
A short phrase for fast exploration, a natural-language one-liner as the everyday default, labeled fields when you want to tune one lever at a time, and timestep prompting when the action has to land on a mark. Start short, then lengthen only where you need more control.
Does Flux 3 generate audio with video?
Yes, and it is on by default. Every clip comes back with audio generated in the same pass, so an impact, footsteps, ambience, or a spoken line is already synced to the action. Name the sound source you want rather than asking for silence, since a request for quiet can collapse into dead air.
How long can a Flux 3 video be?
Flux 3 generates 5 to 20 seconds in a single pass, in whole seconds, or it can pick a duration to suit the prompt. For longer pieces, use video continuation to extend a clip from its final frames, or block several shots inside one generation with hard cuts between them.
How do I keep a character consistent in Flux 3?
Repeat the subject description word for word in every shot of the prompt. Black Forest Labs' schema treats that fixed description as the thing that holds identity across a cut. Keyframes and multi-shot generation then carry the look through the sequence. Combined image and video references are announced as coming next.
How do Flux 3 keyframes work?
You pin stills and Flux 3 interpolates one continuous shot through them. One image sets the opening frame, two fix a start and an end, and up to ten pinned to timestamps become a storyboard the shot has to move through in order, with the model filling in the motion between marks.
How do I use Flux 3?
Write the shot as a short brief naming the subject, its motion, the camera, and one audio cue. Draft it first for a fast preview, judge the direction, then render the keeper at full quality. Extend it with continuation, or pin keyframes when the shot has to hit specific moments.