UGC & Creator Ads

Should you generate B-roll clips before or after lip-sync shots when making an AI UGC ad?

Last updated August 1, 2026

Generate B-roll first. B-roll clips have no voice or mouth movement to match, so you iterate on framing, motion, and product action cheaply before any audio constraint exists. The one exception: lip-sync a single shot early to establish the character's voice model, clone and lock it, then finish all B-roll and generate the remaining lip-sync shots last.

Generate and lock your B-roll-only sequences before dialogue shots because rejecting a B-roll clip costs you nothing but the clip — there is no voice track or mouth movement it has to stay matched to. That matters because most AI-generated clips get discarded: in one documented UGC ad, only 1 of 11 generated clips made the final cut (9% utilization); another used 8 of 20; a localization run rejected roughly 85% of all clips. Run that rejection loop on B-roll, where a regeneration is purely visual, rather than on lip-sync shots where every retry also has to hold the voice.

Sequence the voice deliberately around that order. If your B-roll carries narration, lip-sync one shot first to establish the character's voice model, then ask the invideo agent to clone that voice from the generated clip and lock it for every shot where the character isn't on screen. The invideo agent runs a 2-second audio-similarity comparison before final rendering and auto-regenerates the track if vocal drift between segments exceeds a 1.5% variance threshold, so narration stays consistent across B-roll you generated at different points. Before any dialogue shots, tell the invideo agent to keep the voice identical across every shot — and if the ad has multiple characters, lock a separately generated voice per character rather than sharing one model.

Generate the remaining lip-sync shots last, against the locked voice. Upload the full voiceover file once and tell the invideo agent which line belongs to which shot — it trims the audio on its own, feeds it into Seedance 2.0 with the character reference, and returns the lip-synced clip, so the dialogue passes stay fast even though they carry the tightest constraints.

The working order: 1) one lip-sync shot to establish and lock the voice (skip if your B-roll is silent or caption-and-music only), 2) generate and lock all B-roll sequences, 3) generate the remaining lip-sync dialogue shots against the locked voice, then assemble.

Watch some of these to see what works for you:

Full UGC ad localization workflow showing exactly how to sequence B-roll and lip-sync shots
End-to-end AI UGC ad workflow with lip-sync, B-roll, and real cost breakdown
Real B-roll generation workflow: 19 clips generated, 8 used — see the iteration cost

You can also ask the agent to clone the voice from one of your previously generated shots and use that cloned voice for all shots where the character is not on screen.

— invideo's creative team

Share

More on UGC & Creator Ads