Multi-modal

Gemini Omni Flash 1.1

by Google DeepMind

Google's video model where the edit is a sentence.
Now with 4K output and 40‑second scenes.

Gemini Omni Flash 1.1

Key features

Technical specifications

Gemini Omni

Second Omni Flash release in Google's Omni family

5 tasks

Text, image, reference, edit, and extend

3 to 10s

Per generation; extensions reach 40 seconds total

Up to 4K

360p and 720p native, 1080p and 4K upscaled

10 + 3

Ten images plus three video clips of up to 3s each

Native

Dialogue, effects, and music in the same pass

16:9, 9:16

Landscape and vertical at every resolution

SynthID

Invisible provenance mark on every generated clip

Use cases

Split frame of the same street cafe scene, photoreal on the left and anime on the right

Edits you type

Hand it a clip and say the change: swap the weather, restyle to anime, rewrite the sign. The rest of the frame stays as shot.

A woman speaking to camera mid-sentence in soft window light

Talking characters

Write the line and the character says it, lip-synced, in a voice that survives every cut and every edit that follows.

Reference cards of a bottle, a model, and a location pinned beside the finished perfume ad frame

Reference-true spots

A product still, a presenter photo, and a three-second clip of the move you want become one spot with the product held true.

Four-panel filmstrip of the same cyclist riding through a bakery street, harbor, lighthouse, and coast road

Scenes that grow to 40s

Generate the opening beat, then extend it ten seconds at a time into a 40-second scene with the same cast, setting, and score.

A kinetic sculpture of dominoes tipping a steel marble onto a curved wooden track

Explainers with physics

Gravity, fluids, and collisions follow real-world rules, and Gemini's grounding in science and history keeps the detail right.

An interview frame of a fisherman by a misty harbor with a lower-third reading THE COAST ROAD

Titles rendered in frame

On-screen text comes back readable, from lower thirds to a storefront sign, and a follow-up edit can rewrite what it says.

Prompt examples

Night market

A rain-washed night market, vendors calling, woks sizzling

Edit prompt

Timed beats

Every two seconds, cut to a new rooftop at golden hour

Edit prompt

On-screen text

A diner sign flickers on, letter by letter: OPEN ALL NIGHT

Edit prompt

One continuous shot

A marble run through a sunlit workshop, single unbroken shot

Edit prompt

Spoken line

A climber whispers the summit line into the wind

Edit prompt

Sound-led close-up

Close-up pour-over coffee, every drip and hiss audible

Edit prompt

Simple pricing

Get started for free today, with the option to upgrade or cancel anytime.

Basic

$9/ month
billed as $0 per year

1100 monthly credits

1 user only

All models

Workflows

Standard

$24/ month
billed as $0 per year

3625 monthly credits

1 user only

All models

Workflows

Pro

$45/ month
billed as $0 per year

6350 shared monthly credits

1 user

+ up to 4 more at extra cost

All models

Workflows

Pro Max

$170/ month
billed as $0 per year

24650 shared monthly credits

1 user

+ up to 9 more at extra cost

All models

Workflows

Enterprise

For higher limits

Custom

pricing and billing terms

High-volume credits
Custom seat limits
All models
Workflows
Pricing Gradient

Free

For playing around

$0

forever free

Up to 20 credits
1 user only
Limited models
Workflows

FAQs

What is Gemini Omni Flash 1.1?
Gemini Omni Flash 1.1 is the second release of Google's multimodal video model, shipped in late August 2026. It generates video with synchronized native audio from text, images, and video references, and it edits and extends footage through plain-language instructions. Version 1.1 adds a resolution ladder up to 4K, working video references, first-to-last-frame interpolation, and extensions that take a clip to 40 seconds.
What is new in Gemini Omni Flash 1.1 compared to the first release?
Four things. Resolution is now selectable from 360p to 4K where the first release was fixed at 720p. Video extension is live, growing a clip ten seconds at a time to 40 seconds, where the original capped at a single ten-second generation. Reference videos now actually steer the output, up to three clips of three seconds each. And image-to-video accepts an end frame, so the model can interpolate between two stills you provide.
What resolution does Gemini Omni Flash 1.1 output?
Four options: 360p, 720p, 1080p, and 4K, in 16:9 or 9:16. The model renders natively at 360p and 720p, and produces 1080p and 4K as upscales of that generation. In practice that means drafting cheap at 360p, locking the take at 720p, then requesting the delivery resolution once the shot is right.
How long can a Gemini Omni Flash 1.1 video be?
A single generation runs 3 to 10 seconds, with 8 seconds as the default. From there, extension grows the clip ten seconds at a time up to 40 seconds total, keeping characters, motion, and audio continuous across each join. Extension appends to the end of the clip only; it cannot prepend footage or lengthen the middle.
How does video extension work in Gemini Omni Flash 1.1?
You ask for a continuation in plain language, like "extend this video" or "the scene continues as the camera pans across the mountains". The model reads the last ten seconds of the clip as context and generates the next beat, lightly adjusting the final frames so the join is seamless. It works on clips the model generated and on footage you upload, as long as an uploaded clip is ten seconds or shorter.
What references does Gemini Omni Flash 1.1 accept?
Up to ten reference images and up to three reference videos of three seconds each, alongside the text prompt. Images can pin a character, a product, a setting, or a style, and tags in the prompt assign each one its role. Video references carry a likeness or a motion; any audio inside them is ignored, and asking the model to reason across several full videos at once is not supported.
How does conversational editing work in Gemini Omni Flash 1.1?
Each follow-up prompt edits the previous result instead of generating from scratch, and the model keeps the scene state between turns. Short instructions work best: "make the phone invisible, keep everything else the same" beats a paragraph describing the whole frame. Because every turn builds on the last, one focused change per turn keeps characters and motion stable across a chain of edits.
Does Gemini Omni Flash 1.1 generate audio and dialogue?
Yes, sound is generated in the same pass as the picture: dialogue with lip sync, effects, ambience, and music timed to the action. Write the audio you want into the prompt, because when a prompt says nothing about sound the model invents its own track, and a spoken line is as simple as writing what the character says.
Can I use Gemini Omni Flash 1.1 on Morphic?
Yes. Pick Gemini Omni in the video model picker, write the shot with the sound you want, and generate; the take lands on the Canvas next to results from Veo, Seedance, and Kling for side-by-side comparison. Follow-up prompts edit the same scene, and reference images can ride along with the brief.
Is Gemini Omni Flash 1.1 the same as Veo 4?
No, and the confusion is understandable: both come from Google DeepMind and Omni is often written about as the successor to the Veo line. They are separate models. Veo is the cinematic video family, currently Veo 3.1; Gemini Omni is the multimodal family whose first video model is Omni Flash, now at version 1.1. Google has not released a model called Veo 4, so any "Veo 4" reference you see is either speculation or an informal nickname for Omni.
How is Gemini Omni Flash 1.1 different from Veo 3.1?
Veo 3.1 is tuned for single-shot cinematic fidelity; Omni Flash 1.1 is tuned for iteration. Omni is the one you direct: it edits its own output turn by turn, takes mixed image and video references, extends a take to 40 seconds, and holds a character and voice through the process. For a one-off hero shot Veo still leads on polish; for a scene you refine and grow, Omni Flash is the workflow.