Yes. AI can track text on a moving object two ways: post-production motion tracking, where AI segmentation pins a text layer to a tracked subject after filming, and generation-time tracking — Google Omni Flash can keyframe and track text overlays tied to a moving subject inside the generated clip itself, so the text moves correctly from the first frame.
Post-production tracking is the established route: most modern editors run AI motion tracking that locks a text layer to a subject automatically — you place the text once, the tracker follows the object, no manual keyframes. Community threads confirm this is now a standard, widely available feature rather than a specialist skill.
Generation-time tracking is the newer route, and it works differently: instead of applying text to existing footage, you prompt the model to render the text as part of the video. In Google Omni Flash, you can prompt for keyframed text overlays tied to a moving subject, and the model tracks the text alongside that subject within the generated clip. Across 30+ test outputs, in-frame text tracking alongside moving subjects held up as a genuine capability rather than a fluke — and because the text is rendered inside the scene, its motion, perspective, and lighting match the subject natively. Motion graphics with keyframe animation and consistent, legible text is one of the bigger unlocks in this model class.
The text can carry real data, not just labels. Omni runs on Gemini's internet-scale knowledge layer, which means you can prompt for factually grounded callouts and the model supplies accurate content: in a mock explainer test, it generated a tracked data graphic reading "47% increase in workplace happiness" with the text rendered cleanly and consistently through the motion. For explainer videos, that combines tracking, animation, and factual narration in one generation pass.
Practical constraints to plan around. Omni Flash generates clips at 4, 6, 8, or 10 seconds, so tracked-text sequences longer than that need to be built shot by shot. Output is 720p by default with a free 1080p upscale; 4K costs the equivalent of a full generation — worth it when small tracked text needs to stay legible. As invideo's creative team notes, "Google's Omni model is one of the few models, if not the only model out there that offers native 4K." The tradeoff versus post-production tracking: generation-time text is baked into the pixels, so changing the copy means regenerating the clip, and you're steering with prompts rather than a visual tracking UI.
For explainer-style projects where you want tracked text, data callouts, and narration produced together, the invideo agent can generate the full video — script, visuals, on-screen text — from a single brief, so you don't have to assemble the tracking pass separately.
Watch some of these to see what works for you:
Google's Omni model is one of the few models, if not the only model out there that offers native 4K. The moment we have a 4K model that has very very very strong VFX capabilities, we will finally have an AI model that is ready for big screen primetime.
— invideo's creative team