Why do AI-generated ads sometimes reproduce captions from a reference video?
Last updated August 1, 2026
AI-generated ads reproduce captions because video models treat everything in an attached reference frame as visual content to recreate — including burned-in text. The model can't distinguish an overlay caption from the scene itself, so it renders the old captions into new footage. The fix: use the reference for analysis only, and generate from clean assets.
Video generation models read a reference video pixel-by-pixel, and burned-in captions are just pixels. When you attach a reference ad directly to a generation prompt, the model interprets the caption text as part of the scene it's being asked to reproduce — the same way it reproduces the lighting, framing, and set. It has no signal that the text was an editorial overlay added after the shoot rather than something physically present in the shot, so old captions — often in the wrong language for your new market — get baked into new clips.
This shows up most often in ad recreation and localization workflows, where the whole point is to feed the AI a winning reference ad. The prevention is a separation of roles: let the AI analyze the reference, but never attach it to the actual generation prompt. In one documented localization run inside the invideo agent, the team had the agent watch the reference ad, detect every cut (9 cuts in that ad), and extract one representative frame per scene to build a shot-flow breakdown — then deliberately excluded the reference video from all generation prompts. Generation was driven only by the pre-built character sheet, location sheet, and shot breakdown, so no caption text existed anywhere in the inputs. That workflow held 100% product and text consistency across localized ads tested on 3 different ads in 2 completely different markets.
If your reference has captions and you can't work from extracted assets alone, use a clean version of the ad — a pre-caption export or a stripped copy — as the generation reference instead. Then regenerate on-screen text deliberately as a separate step: in the same documented workflow, localized CTA copy (e.g., Japanese '初回購入は最大45%オフ' — up to 45% off on first purchase) was written fresh for the target market rather than inherited from the source video, which is also how the copy gets culturally adapted instead of literally copied.
One adjacent point: even with clean references, on-screen text in AI video can artifact — that's a model-level limitation of video generation models like Seedance 2.0, not a workflow failure. Keep critical text as a separate deliberate pass (regenerated screens or supers added in the edit) rather than trusting the model to render it inside the footage.
Watch some of these to see what works for you:
I made sure that the prompt didn't actually have the original reference video attached. The reference ad, if you remember, had captions on it. So, if the agent attached the ad while prompting, the AI model would pick up those captions and it would put them into the new generations, which you don't want.
— invideo's creative team, documenting an ad localization production