Gemini Omni Flash

Create videos, avatars, and explainers on Gemini Omni Flash, the model that holds character, physics, and scene memory through every edit. Invideo agent runs the conversation with the model so you never have to write the instructions.

Trusted by teams at
Phantom X
Wonder Studios
Google
NVIDIA
Salesforce
Visa
IBM
monday.com
Cisco
Mastercard
Meta
LinkedIn
Netflix
Ford
Hilton
Siemens
HubSpot
Phantom X
Wonder Studios
Google
NVIDIA
Salesforce
Visa
IBM
monday.com
Cisco
Mastercard
Meta
LinkedIn
Netflix
Ford
Hilton
Siemens
HubSpot

Why serious creatives choose Gemini Omni Flash

Editing that only touches what you named

Say the background should be a rainy street, swap the product, relight the scene, remove the object in the corner. Omni Flash changes that and only that, and carries the rest forward untouched.

Any input, in any combination

Gemini Omni Flash takes a wide variety of inputs: text, images, audio, video, sketches. A character photo, a location shot, a voice reference, and a written beat go in together, and the model reads them as one instruction rather than four competing ones.

Physics it actually understands

Gravity pulls, things collide with real weight, and liquids move like liquids. That understanding is built into the model, so a chain reaction resolves instead of falling apart halfway through.

Grounded in what Gemini knows

Omni Flash draws on Gemini’s world knowledge, so it connects language, imagery, and meaning instead of pattern-matching its way to something adjacent. Ask for a protein fold or an engine cycle and the science stays honest.

Characters that survive the edit

Faces, wardrobe, lighting, and scene continuity hold across every turn of the conversation. Where other models drift a little further from your character with each revision, this one keeps them.

Avatars that look and sound like you

Record your voice and likeness once and generate videos that carry your face, your expressions, and your voice from a text prompt. The presenter in shot one is the presenter in shot forty.

Text that tracks with the shot

Attach text to a moving object, a person, or a surface, and it holds position, perspective, and style as the shot moves. Signage, labels, and callouts stay stuck to the thing they belong to.

Motion and style, taken from reference

Hand it a movement you like and it transfers that motion to another character or scene. Same for a visual style: point at the look, and the output adopts it without turning into a filter.

How Gemini Omni Flash works with invideo agents

On invideo, Gemini Omni Flash runs inside an agentic workflow: the agent routes work to it where it wins, writes its prompts, attaches the right references, and keeps your characters and brand rules on every generation. Here is what that looks like in practice.

The agent picks the right model intelligently.

Omni Flash sits on a roster of 200+ models, and the agent routes work to it where it wins: fast iteration, conversational shaping, explainers that need the science right, avatar work, and physical action. You never choose a model unless you want to, and if you have routing rules of your own, they hold for the whole project.

You direct, the agent writes the prompts.

Omni Flash reads mixed input better than almost anything on the roster, which means what you attach and how you frame it decides the output. You describe the shot in plain language, and the agent assembles the brief: the right references, in the right order, written the way the model reads best.

Your production rides along on every generation.

Omni Flash remembers your scene inside a conversation. The agent remembers your project across all of them: your characters, your avatar, your brand rules, your look, attached to every generation and carried into every new thread, so nothing gets re-explained at scene forty.

Talk it into shape here, finish it anywhere.

Omni Flash is the fastest way to find a shot, one instruction at a time, until it is right. That found shot does not have to stay here: the agent can take the direction you landed on and run it through a model built for the final render, so you iterate cheap and deliver finished.

You always stay in control.

You set how much the agent does on its own: let it generate everything, ask before videos, or see every prompt before it generates. That setting is yours to make, and yours to change.


Who is Gemini Omni Flash for?

Gemini Omni Flash supports conversational, multi-input video generation across filmmaking, advertising, and education workflows. With invideo agents underneath it, each team gets the model without learning the machinery.

Filmmakers

Create shots from an idea alone, or bring what you have: a character photo, a location plate, a voice reference, a scribbled frame. Omni Flash helps a scene get found through conversation rather than guessed at through prompts.

Performance Ads and UGC teams

Create hooks, product moments, and avatar-led variations, then adjust each one by saying what should change. Useful for testing cycles where the winner needs twenty versions and none of them should look regenerated.

Brand and product marketers

Create ads that respect brand colors, product shape, and on-screen text, from one product photo and a brief. Useful when the product has to stay exactly itself while everything around it changes.

Educators and course creators

Create explainers where the physics and the science hold up: chain reactions, protein folds, engine cycles, historical scenes. Omni Flash helps a concept get visualized correctly, not just attractively.

Content creators

Create social videos and avatar-led posts at the pace a feed demands, refining each one in conversation until it lands. Useful when the idea and the finished post should happen in the same sitting.

Helping creatives stay creative

Multiplayer mode

Collaborate in real time with live cursors to show what everyone's working on.

RRebeccajust now
Can we push the train arrival 2 sec later? Feels rushed.
AAiko2m
Love the warmth here — keep this lighting for the reunion shot.

Storyboarding

Turn any script or idea into a shot-by-shot plan, then tweak as needed before generating.

Script writing

Write your script inside invideo, and ask an AI co-writer for help if you'd like.

Timeline editor

Picture Premiere Pro with full AI.

Build your own agents

Create custom agents to fill specific roles like cinematographer, music designer, and more.

From solo creatives to creative enterprises

Private & SecureSOC 2 & ISO AlignedGDPR Compliant

World-class investors stand behind invideo.

Backed by the firms behind Stripe, Spotify, Flipkart, and ByteDance.

Pricing

All paid plans include:

Access to 200+ image, video, audio, music models including Seedance 2.0, Veo 3.1, Kling 3.0, Nano banana pro & Elevenlabs music.

Access to top stock providers like iStock, Storyblocks & more.

Model & agent prices are subject to change.

On-demand credit top-ups available.

Gemini Omni Flash FAQs

What is Gemini Omni Flash?

Gemini Omni Flash is the first model in Google’s Gemini Omni family, announced at Google I/O in May 2026. It takes any combination of text, images, audio, and video and generates video with native audio, grounded in Gemini’s real-world knowledge, and it edits through conversation rather than re-prompting. On invideo, it runs inside the agent’s roster: you direct in plain language and the agent writes the instructions.

What is the maximum duration for a video I can create in Gemini Omni Flash?

Clips run up to 10 seconds. Google has described that as a rollout decision rather than a limit of the model, so it may move.

Where can I use Gemini Omni Flash?

On invideo, open Agents & Models and select Gemini Omni Flash, then upload text, images, video, or audio as reference and describe what you want. Google also ships it in the Gemini app, Google Flow, and YouTube Shorts.

Is Gemini Omni better than Sora, Kling or Runway AI video generators?

Better at different things. Omni Flash leads on conversational editing, mixed input, and physics grounded in real understanding. Others lead on resolution, clip length, or peak single-prompt render quality. On invideo you do not have to pick: all of them sit on the same roster, and the agent routes each shot to whichever wins it.

Will Gemini Omni be replacing Veo 3.1?

No. Google runs them as separate lines: Omni is the any-to-any conversational model, Veo remains the realism specialist. On invideo, both stay on the roster and the agent picks per shot, so an Omni conversation and a Veo render can serve the same scene.

Do I need to learn Gemini Omni Flash’s prompting to use it on invideo?

No. Omni Flash rewards well-assembled mixed input, the right references framed the right way, and the agent handles that for you. You describe the shot the way you would to a person, and the agent builds the brief.