Gemini Omni Flash is Google's lightweight multimodal video generation model: it takes text, image, audio, or video inputs and generates 4-, 6-, 8-, or 10-second clips at 720p, with free 1080p upscale and native 4K. It is currently the only AI video model with native avatar generation, plus in-paint, cleanup, swaps, motion graphics, style remixing, and in-frame text tracking.
Gemini Omni Flash is Google's lightweight video generation model, launched at Google I/O in May 2026, that accepts text, image, audio, and video as inputs and supports conversational, natural-language editing of its outputs. The capability breakdown below comes from a benchmark by invideo's creative team across 30+ test outputs.
Native avatars — its unique feature. Omni Flash is currently the only AI video model offering native avatar generation: a short face calibration plus a voice session builds a personal avatar you can drop into generated video. The voice setup is minimal — you read double-digit numbers to the camera, no sentences or paragraphs, and the model infers intonation, pauses, and syllable handling from that alone. In testing, voice replication came out stronger than facial replication, and lip sync holds consistently up to a 6–7 second ceiling with a single speaker in frame. You can also inject a real person as a consistent character using just two static reference images, no live footage needed.
Generation specs. Clip lengths are fixed at 4, 6, 8, or 10 seconds. Output defaults to 720p, upscales to 1080p at no cost, and native 4K costs the equivalent of a full generation. That native 4K option is what no other current AI video model offers — the missing piece for cinema-grade output is texture quality, not resolution.
Editing and compositing toolset. Beyond generation, Omni Flash works on existing footage: in-paint and cleanup insert or remove objects from a clip; the swap feature replaces backgrounds, environments, or clothing around a subject while preserving the subject's roto and edges; and in-frame text tracking lets you prompt keyframed text overlays that stay locked to a moving subject. Motion graphics with keyframe animation and consistent text is a genuine unlock — one test produced a mock explainer with an accurate on-screen "47% increase in workplace happiness" data visualization. It also handles multi-shot clips and style remixing of existing footage.
The Gemini intelligence layer. Because the model sits on Gemini's internet-scale knowledge base, you can prompt factually grounded explainer videos — real anatomical and scientific facts narrated over generated motion graphics, without feeding the facts in yourself.
How it benchmarks against Veo 3.1. Visual texture and lighting are a generational step up from Veo 3.1, though still in the same visual family. Time code accuracy when prompting specific moments is sharper in Omni. Physics adherence is strong for ordinary prompts; complex physics prompts land roughly 50/50, but the successful ones are exceptional. One capability gap runs the other way: scene extension currently only works on clips generated in Veo 3.1, not Omni-generated content.
Known limits. Omni will not generate real contact-based actions or anything violence-adjacent — consistent across all Veo-family models. Camera angle prompting is hit-or-miss, and failed angle changes tend to distort scene geography rather than just the angle. Multiple people speaking in the same frame remains weak — a cross-model problem, and splitting the frame into separately generated halves and compositing them is a workaround, not a real capability. Stop motion outputs oscillated between 12 FPS and 8 FPS, and current textures are not yet at prime-time cinema standard.
Availability and pricing. Omni Flash ships through the Gemini app, Flow, and YouTube Shorts, with a preview developer API priced at $0.10 per second of generated video.
Watch some of these to see what works for you:
Google's Omni model is one of the few models, if not the only model out there that offers native 4K. The moment we have a 4K model that has very very very strong VFX capabilities, we will finally have an AI model that is ready for big screen primetime.
— invideo's creative team