AI video generators mimic physics visually without modeling physical laws — plausibility, not simulation. In a 30+ output test of Google Omni Flash, non-contact physics (fluids, lighting, environmental motion) held up strongly, while complex physics prompts landed roughly 50/50, and contact-based actions failed or were refused outright — a pattern consistent across Veo-family models.
Current models predict what a physically plausible next frame looks like; they do not compute forces, mass, or collisions. Academic benchmarks like Physics-IQ confirm the gap: visual realism is no guarantee of physical consistency, and models score well on appearance while failing conservation-of-momentum-style tests (arXiv).
Where accuracy is strong. Non-contact dynamics are the reliable zone. Across 30+ test outputs on Google Omni Flash, physics adherence outside violence-adjacent prompts was strong — fluid motion, cloth, environmental behavior, and light interaction track convincingly. Time-code-specific prompting was also sharper in Omni than in Veo 3.1, which matters when you need a physical event to land at an exact moment in the clip.
Where it becomes a coin flip. Complex physics prompts — multiple interacting objects, chained cause-and-effect — ran approximately 50/50 in the same testing. The failure mode isn't subtle wobble; when a generation misses, it misses structurally. Camera-angle changes show the same fragility: failed angle prompts tend to distort the scene's geography rather than just the angle, breaking spatial consistency across the shot.
Where models refuse or break entirely. Real contact-based actions — impacts, grappling, anything violence-adjacent — will not generate correctly in Omni, a limitation consistent across all Veo models. Community testing echoes this: car collisions, water displacement, and destruction physics are the most-reported failure categories across models (Reddit discussion), though early community reports on Seedance 2.0 point to meaningfully improved destruction and contact physics (Reddit).
How to work with a 50% ceiling. The successful half of complex-physics generations is genuinely usable — when the model gets it right, it gets it really right. So treat physics-heavy shots as a multi-generation job: run several takes of the same prompt, keep the take where the physics holds, discard the rest. Model choice matters here too — inside invideo, all current models (Veo, Kling, Seedance 2.0, Runway) are available, and the invideo agent routes each shot to the model best suited to it, so a contact-heavy shot and a fluid-motion shot don't have to run through the same model. On the research side, physics-aware architectures that put a simulator or explicit physical reasoning in the generation loop (e.g. DiffPhy at Johns Hopkins) are the direction the accuracy problem is being attacked from — meaning today's 50/50 on complex prompts is a moving number, not a fixed one.
Watch some of these to see what works for you:
this is kind of 50/50. It gets it right sometimes, it gets it wrong sometimes, but what it gets right, it gets it really right.
— invideo's creative team, on Omni's complex physics prompt handling