How does Google Gemini Omni avatar generation work — and what are the requirements?
Last updated August 1, 2026
Gemini Omni generates a personal avatar from a two-part calibration: a head-movement face scan plus a voice pass where you read double-digit numbers aloud — no sentences or paragraphs required. Requirements: you must be 18+, in a supported region, and on a paid Google AI plan; you then summon your avatar in prompts using @username syntax.
Set up the avatar in two calibration steps, then call it inside your prompts.
Step 1 — face calibration. Omni runs a brief onboarding scan where you move your head on camera so the model captures your face from multiple angles. In testing, facial replication came out solid but not perfect — voice replication is the stronger half of the clone. As invideo's creative team put it after running the setup: "What came out the other end was actually a decent replica of my face but an incredible replica of my voice."
Step 2 — voice calibration by reading numbers. You read double-digit numbers to the camera — that's the entire voice sample. "It was just double digit numbers that I was saying to the camera — it didn't ask me to say a sentence, read out a paragraph." From numbers alone, the model infers your intonation, pauses, and syllable handling, which is why the voice replica lands so accurately with so little input.
Step 3 — summon the avatar in prompts. Once calibrated, reference your avatar with @username syntax in a prompt, describe the scene and the lines, and Omni generates the video with your face and voice performing it. Omni is currently the only AI video model offering native avatar generation, so no external cloning tool sits in the pipeline.
Requirements. Per Google's documentation: you must be 18 or older, located in a supported region, and subscribed to a paid Google AI plan. Every avatar video carries an embedded SynthID watermark for provenance, so outputs are verifiable as AI-generated.
Output specs and practical limits. Clips generate at 4, 6, 8, or 10 seconds. Output defaults to 720p, upscales to 1080p at no cost, and 4K costs the equivalent of a full generation. Across 30+ tested outputs, 6–7 seconds proved the ceiling for consistent lip sync with a single speaker in frame — keep talking segments under that and cut between clips for longer monologues. Keep one speaker per frame: multiple people talking in the same shot remains a weakness across all current AI video models, and compositing two half-frames separately is a workaround, not a capability.
If you don't want the full calibration: you can instead upload two static reference images of a real person to generate a consistent character appearance in Omni video — no live footage or scan needed — though this gives you visual consistency only, not the cloned voice.
Watch some of these to see what works for you:
It was just double digit numbers that I was saying to the camera — it didn't ask me to say a sentence, read out a paragraph.
— invideo's creative team, on the minimal input Omni's voice cloning requires