Why should you generate and lock B-roll shots before lip-sync shots in a localized video ad?
Last updated August 1, 2026
Lock B-roll before lip-sync shots because B-roll has no voice or mouth movement to match — you evaluate visuals only, so each iteration is cheaper and faster. Documented localization runs rejected ~85% of generated clips, so sequencing the no-audio shots first isolates the hardest variable — lip-sync against a translated voiceover — until the visual foundation is locked.
Generate and lock every B-roll-only sequence first because those clips carry one evaluation criterion instead of three: you judge character, location, and product consistency without also checking whether mouth movement matches a translated voiceover. A lip-sync shot that fails on any one axis — visuals, voice, or viseme timing — gets regenerated whole, so every variable you remove from a generation pass directly cuts rejected clips.
That matters because localization iteration overhead is real: in one documented run, roughly 85% of all generated clips were rejected, and each localized ad required regenerating around 13 video clips plus 4 reference sheets, a translated voiceover, and 5 translated app UI screens. Across documented runs, localized ads landed at $70–$145 per ad ($425 for six ads across two markets in one run; 570 credits per ad in another), with ~2.5 hours to recreate each ad — and those figures already include the rejects. Burning that rejection rate on shots that also have to hit lip-sync is the expensive version of the same work. invideo is an agentic video creation tool, and inside the invideo agent this sequencing runs as a lock-and-regenerate loop: lock the B-roll clips that pass, regenerate only what fails, then move to dialogue.
Locked B-roll also becomes the consistency anchor for the harder shots that follow. With the visual world fixed — character look, location lighting, product scale — the lip-sync shots only have to solve audio-visual matching, not re-establish the ad's look. On the audio side, the order pays off twice: before generating any dialogue shots, tell the invideo agent to keep the voice identical across every shot, and for B-roll narration, clone the voice from a locked lip-synced clip rather than generating a new one. The invideo agent runs a 2-second audio-similarity comparison before final render and auto-regenerates the track if vocal drift exceeds a 1.5% variance threshold — a check that only works if the B-roll visuals are already locked and waiting for narration.
Finally, the sequencing compounds across markets. The edit structure, music bed, and 7 shot beats stay constant in localization — only cast, location, voiceover, and UI text change — so a clean B-roll-first pass in your first market becomes the repeatable template; documented runs hit 6 localizations in a day once the first was locked. If you want the same discipline one level earlier, iterate framing on still images before spending video credits at all.
Watch some of these to see what works for you:
You can also ask the agent to clone the voice from one of your previously generated shots and use that cloned voice for all shots where the character is not on screen.
— invideo's creative team