AI Ads

How do you prevent burned-in captions from appearing in AI-generated ad variations?

Last updated August 1, 2026

Burned-in captions reappear because the caption-contaminated reference video is attached to generation prompts — the model copies the old text into every new clip. Prevent it by excluding the reference from generation and directing the invideo agent with clean proxy assets instead: a character sheet, a location sheet, and a shot-flow breakdown. Add each variation's captions as an editable layer in the edit.

Separate the reference ad's two jobs: it can inform the plan, but it must never be attached to a generation prompt. invideo is an agentic video creation tool where the invideo agent can analyze a reference ad, build proxy assets from it, and route each shot to the right model — which is exactly the mechanism that keeps old captions out.

1. Use the reference for analysis only. Upload the winning ad once so the invideo agent can deconstruct it — in one documented localization run it detected all 9 cuts automatically and preserved 7 shot beats in the recreation. The output you keep is the shot-flow breakdown (shot order, duration, action, framing per cut), not the video itself.

2. Build clean proxy assets to drive generation. Have the invideo agent generate a character sheet and a location sheet — GPT-Image-2 works well for location sheets because it handles realistic environments and accepts reference attachments, at roughly 5 minutes per sheet. These caption-free assets, plus the shot-flow breakdown, carry everything the video model needs to recreate each scene.

3. Exclude the reference video from every generation prompt. This is the actual prevention step: if the contaminated video rides along as an attachment, the model reproduces its captions in the new footage. Verify it — in invideo you can click any generation to see the exact prompt and attachments used, so you can confirm the reference video isn't riding along before you commit a batch. A team that ran this discipline across 3 ads in 2 markets hit 100% product and text consistency at about $70 and 2.5 hours per localized ad.

4. Add captions as an editable layer, not baked pixels. Generate clean plates, then lay supers and captions on top during assembly — in Slate inside invideo or in Premiere Pro. Each market variation then gets its own translated caption layer without touching the footage, which is also how you avoid re-generating clips just to change copy.

5. If you must attach the reference, strip it clean first. Some complex hooks need the actual video for timing and framing that text can't convey. In that case run the reference through an AI inpainting-based caption removal pass (or crop the caption band out) before attaching it — removal tools work but can leave artifacts over busy backgrounds, so inspect the cleaned reference before it enters your prompts.

Watch some of these to see what works for you:

Full guide: localizing UGC ads without burned-in captions contaminating new footage

I made sure that the prompt didn't actually have the original reference video attached. The reference ad, if you remember, had captions on it. So, if the agent attached the ad while prompting, the AI model would pick up those captions and it would put them into the new generations, which you don't want.

— invideo's creative team

Share

More on AI Ads