AI VFX

What is in-frame text tracking in AI video and how does it work?

Last updated August 1, 2026

In-frame text tracking is an AI video capability where text overlays are keyframe-animated and spatially locked to a moving subject inside the generated clip itself — the text moves with the person or object rather than sitting as a static burn-in. You get it by prompting the model (currently Google Omni Flash) to track text to the subject at generation time.

In-frame text tracking means the model generates text as part of the scene: a label, caption, or stat that stays pinned to a moving subject — following them as they walk, turn, or as the camera moves — with the tracking baked in at generation time. This is different from traditional motion-graphics workflows, where you generate or shoot the footage first, then use a motion tracker in an editor to pin text to a subject in post. With in-frame tracking, you skip the tracking pass entirely: the model handles the keyframing as it renders the frames.

How it works in practice. You describe the text and its behavior in the prompt: what the text says, which subject it should attach to, and how it should move — for example, a stat callout that follows a presenter across the frame. The model keyframes the text against the subject's motion internally, so the overlay holds position relative to the subject rather than the frame. In testing across 30+ outputs on Google Omni Flash, this worked alongside its broader motion-graphics capability — keyframe animation with consistent, legible text across the clip, which has historically been a weak point for AI video models. Text tracking on moving subjects was flagged as one of the meaningful new unlocks in that testing.

Factual grounding comes from the intelligence layer. Because Omni sits on Gemini's internet-scale knowledge, the text it generates can be factually accurate, not just typographically stable. In one mock explainer test, the model rendered a data callout — a "47% increase in workplace happiness" stat — as an accurate, tracked motion graphic inside the video. That combination of correct data and stable in-frame text is what makes the feature usable for explainer and product content rather than just decoration.

When to use it vs post-production tracking. Use in-frame generation when the footage itself is AI-generated and the text is part of the concept — explainers, stat callouts, animated labels. Post-production tracking still makes sense when you need frame-accurate control over typography, brand fonts, or revisions after the fact, since generated text can't be edited without regenerating the clip. Practical parameters if you're generating with Omni Flash: clips come in 4, 6, 8, or 10-second lengths, output defaults to 720p with a free 1080p upscale, and 4K costs the equivalent of a full generation — relevant if your tracked text needs to stay sharp at delivery size.

Model choice matters here. Not every video model handles text tracking — Veo 3.1, Kling, and Seedance 2.0 each have different strengths, and text consistency is where Omni currently leads. Inside invideo you have access to all of these models, and the invideo agent routes each shot to the right one — so a motion-graphics shot with tracked text goes to the model that can actually hold the type stable.

Watch some of these to see what works for you:

See in-frame text tracking and motion graphics tested live on Veo Omni Flash

Google's Omni model is one of the few models, if not the only model out there that offers native 4K. The moment we have a 4K model that has very very very strong VFX capabilities, we will finally have an AI model that is ready for big screen primetime.

— invideo's creative team

Share

More on AI VFX