AI lip sync stays accurate for about 6–7 seconds with a single speaker in frame — the consistent ceiling found across 30+ test outputs of Google Omni Flash. Past that, mouth-to-audio alignment drifts. Multiple speakers in one frame break sync far sooner, and post-generation lip-sync tools applied to real footage hold minutes, not seconds.
Treat 6–7 seconds as your working limit for a single speaker in natively generated AI video. Across 30+ test outputs of Google Omni Flash, that was the point where lip sync stopped performing consistently — mouth shapes and audio stay locked below it and start drifting above it. This applies to native generation, where the model produces the face and the dialogue together in one pass.
Plan your dialogue coverage around that number. Omni Flash generates clips at 4, 6, 8, or 10 seconds — keep talking shots at 4 or 6 seconds and chain beats together rather than pushing one 10-second take, since the last seconds of a long take are where sync degrades first. Note that scene extension currently only works on clips generated in Veo 3.1, not Omni-generated content, so structure dialogue as separate short shots from the start instead of planning to extend one. Working inside invideo, you can have the invideo agent break a dialogue scene into shot-length beats and cut between them, keeping every individual clip under the ceiling.
The 6–7 second figure assumes one speaker. Two or more people talking in the same frame is a persistent weakness across every AI video model tested — not an Omni-specific flaw — and sync reliability collapses well before the single-speaker ceiling. Generating each half of the frame separately and compositing the speakers together is a workaround, not a model capability; cut between single-speaker shots instead, which also keeps each clip inside the accurate range.
The duration answer changes by tool category. Native generation holds seconds; audio-driven lip-sync tools that map new audio onto existing recorded footage hold far longer — one documented user report puts breakdown around the 6:50 mark of runtime. So if you see multi-minute accuracy claims, they refer to re-syncing real footage, not generating a speaking character from scratch.
The underlying constraint is temporal stability, not lips specifically — the same testing found stop-motion outputs oscillating between 12 FPS and 8 FPS within a clip. Any output element that must stay coherent frame-to-frame over time, lip sync included, gets less reliable the longer a single generation runs.
Watch some of these to see what works for you:
In our research we found that 6 or 7 seconds is kind of the ceiling of where the model is going to perform consistently well.
— invideo's creative team