Video generation

MiniMax H3 Max

by fal.ai

fal's post‑trained MiniMax H3.
Tuned for prompt adherence, rebuilt for speed.

MiniMax H3 Max

Key features

Technical specifications

MiniMax H3

Post-trained by fal from H3's open weights

2 modes

Text to video and image to video, both live

5 to 15s

Five to ten seconds is the recommended range

Up to 768p

Native 480p and 768p, a step above 720p HD

24 FPS

The cadence film and broadcast are shot at

Native

Stereo sound generated alongside the picture

11 languages

Inherited from H3, which is stable across 11

6 ratios

21:9, 16:9, 4:3, 1:1, 3:4, and 9:16 supported

Use cases

Kinetic typography

A drawn trace resolves into legible letterforms and holds. Type stays correctly spelled in Latin and in Japanese.

Isometric 3D builds

Tiles rise, towers extrude and pathways connect, so a product tour or a data story assembles itself on screen.

Technical schematics

Line drawings that build themselves: a plan lifts into isometric, then dimension callouts and annotations fade up.

Collage and cut-out

Torn-edge fragments and printed halftone textures slide in and lock into a subject, a look that reads as handmade.

Print and riso looks

Two-colour ink, coarse dot texture and deliberate misregistration, with the shift of a second printing pass.

Sound-led scenes

Picture and audio are predicted together, so a scene built around one precise sound arrives already carrying it.

Prompt examples

Kitchen morning

A kettle whistles as steam fogs a kitchen window

Edit prompt

Mountain ledge

Two climbers check knots on a windswept ledge

Edit prompt

Tailor's bench

A tailor pins a hem under a bright work lamp

Edit prompt

Turning tide

Waves flatten a sandcastle as the tide turns

Edit prompt

Commute

A cyclist threads morning traffic in low sun

Edit prompt

Harvest

Dust rises as a combine cuts the last wheat row

Edit prompt

Simple pricing

Get started for free today, with the option to upgrade or cancel anytime.

Basic

$9/ month
billed as $0 per year

900 monthly credits

1 user only

All models

Workflows

Standard

$24/ month
billed as $0 per year

3200 monthly credits

1 user only

All models

Workflows

Pro

$45/ month
billed as $0 per year

6200 shared monthly credits

1 user

+ up to 4 more at extra cost

All models

Workflows

Pro Max

$170/ month
billed as $0 per year

24000 shared monthly credits

1 user

+ up to 9 more at extra cost

All models

Workflows

Enterprise

For higher limits

Custom

pricing and billing terms

High-volume credits
Custom seat limits
All models
Workflows
Pricing Gradient

Free

For playing around

$0

forever free

Up to 20 credits
1 user only
Limited models
Workflows

FAQs

What is MiniMax H3 Max?
MiniMax H3 Max is a post-trained version of MiniMax H3, built by fal rather than by MiniMax. fal took H3's open weights, retrained them on their own data for stronger prompt adherence and better aesthetics, then served the result on their own hardware. It returns five to fifteen second clips at up to 768p with native stereo audio.
Is MiniMax H3 Max made by MiniMax?
No, and the name misleads. MiniMax built and open-sourced H3. fal post-trained those weights into the variant called H3 Max and hosts it exclusively. MiniMax has never announced a model by that name, so treat H3 Max as fal's tuned edition of an open MiniMax model, not a MiniMax release or an official H3 tier.
How is MiniMax H3 Max different from MiniMax H3?
Three ways. H3 Max is post-trained for prompt adherence and aesthetics on top of the base weights. It is tuned for speed rather than maximum resolution, stopping at 768p where H3 reaches 2K. And it covers text to video and image to video only, while H3 adds reference generation from mixed images, clips, and audio.
How fast is MiniMax H3 Max?
fal targets sub-three-second text to video on a five-second 768p clip, against the 30 seconds to several minutes video normally takes. It runs on fal's multi-node inference engine, several machines working on one request together. Keep prompt expansion on balanced: the slower quality setting adds about thirty seconds and gives that speed advantage straight back.
Does MiniMax H3 Max do image to video as well as text to video?
Both are live. Text to video turns a written scene into a clip. Image to video animates a still you supply, and the same mode accepts a first and last frame, so the motion runs between two images you picked. A third mode, reference to video, holding a character consistent, is announced as coming later.
Why does MiniMax H3 Max stop at 768p?
Because 768p is what the open weights generate. H3 is a three-part system: a preprocessor that interprets the brief, a base model that renders at 768p, and a separate stage that regenerates it at 2K. MiniMax open-sourced only the base model, so anything post-trained from those weights has no 2K stage to hand off to.
Does MiniMax H3 Max generate sound?
Yes. The underlying model predicts video and audio together rather than scoring the clip afterwards, so a generation returns stereo sound with the picture. That makes audio worth writing into the brief: name the effects you want, the room tone, and the exact moment a cue should land, rather than arranging all of it later.
Can I use MiniMax H3 Max on Morphic?
Yes. Pick MiniMax H3 Max in the model picker, switch the prompt bar to video, and write the shot as a brief covering subject, action, camera, light, and the sound you want. Takes land on the Canvas, so you can run the same brief through another model and compare them side by side before choosing.