AI VFX

How do you make text follow a moving subject in an AI-generated video?

Last updated August 1, 2026

Prompt the tracking into the generation itself. Google's Omni model supports in-frame text tracking: write the exact on-screen text into your prompt and name the subject it anchors to, and the model keyframes the overlay to follow that subject across the clip — no separate motion-tracking or compositing pass. For clips from other models, track the text in post.

Write the text into the generation prompt, verbatim and in quotation marks, then describe the anchoring behavior explicitly: name the subject, where the text sits relative to it, and that it stays locked as the subject moves. A working pattern: "A cyclist rides through a city street; a floating label reading 'Heart rate: 142 bpm' stays locked above her helmet, tracking her position for the full clip." The more specific the anchor point (above the head, on the jersey, beside the left shoulder), the more consistent the tracking.

This works because motion graphics with keyframe animation and text consistency is a major unlock in Omni — the model handles the frame-by-frame keyframing internally instead of leaving it to your editor. In-frame text tracking alongside moving subjects held up across 30+ test outputs, and because Omni runs on Gemini's intelligence layer, data callouts stay factually coherent: one mock explainer test rendered a tracked "47% increase in workplace happiness" stat accurately as an in-frame data visualization. If your tracked text carries real figures or scientific claims, you can lean on that layer rather than hand-checking every label.

Plan the tracked moment inside one generation. Omni clips come in 4, 6, 8, or 10 second lengths, and extend currently only works on Veo 3.1-generated clips — so a text overlay that must follow a subject continuously has to complete its move within a single generation. For delivery quality: output defaults to 720p, the 1080p upscale is free, and 4K costs the equivalent of a full generation — worth budgeting when the text needs to stay crisp at large sizes. As Hridaye, invideo's creative director, puts it: "Google's Omni model is one of the few models, if not the only model out there that offers native 4K. The moment we have a 4K model that has very very very strong VFX capabilities, we will finally have an AI model that is ready for big screen primetime."

On model choice: this is currently an Omni-native capability — Veo 3.1 doesn't keyframe text to subjects in-generation, and Omni also prompts time codes more sharply than Veo 3.1, which helps when you want text to appear or dismiss at a specific second. Working inside invideo, which carries all the current generation models, you can describe the tracked-text behavior in your brief and the invideo agent routes the shot to a model that handles it natively rather than you testing models one by one.

If you already have a generated clip from a model without native text tracking, the fallback is the classic post route: motion-track a point on the subject in your editor and pin the text layer to it — the standard matchmoving approach editors have used for years, and the usual answer in editing communities for tracking text over a moving person. It works on any footage, but it's a manual pass per shot, which is exactly what prompting the tracking into the generation removes.

Watch some of these to see what works for you:

See how the invideo agent handles in-frame text tracking with Omni

Google's Omni model is one of the few models, if not the only model out there that offers native 4K. The moment we have a 4K model that has very very very strong VFX capabilities, we will finally have an AI model that is ready for big screen primetime.

— Hridaye, invideo's creative director

Share

More on AI VFX