How do you auto-trim a voiceover to match lip-sync video shots using AI?
Last updated August 1, 2026
Upload the full voiceover file once, then tell the invideo agent which line belongs to which shot. The agent trims the audio per shot autonomously, feeds each segment plus the locked character reference into Seedance 2.0, and returns a lip-synced clip — no manual DAW cuts, no re-uploading isolated lines.
Skip pre-cutting the VO in an editor. The invideo agent is an agentic video tool that holds your project context (script, character sheet, shot list) and routes generation across models like Seedance 2.0, Kling, and Veo — so it already knows which line of your script maps to which shot. Give it the whole voiceover and the mapping, and it does the rest.
Step 1 — Lock the shot breakdown and character first. Before any audio work, have a locked shot list (shot number, duration, dialogue line) and a locked character sheet with version numbers referenced. The agent uses the dialogue field on each shot as the lookup key when it trims. Skip this and the agent has nothing to match audio against.
Step 2 — Upload the full VO file once. Drop the complete voiceover (generated in the agent via ElevenLabs, or your own recording) into the chat. Do not pre-trim. One file, one upload — it sits in project context and gets reused across every lip-sync shot in the ad.
Step 3 — Tell the agent which line belongs to which shot, in plain English. "This is the full VO. For shot 3, use the line 'I tried it for two weeks.' For shot 5, use 'and the energy is real.'" The agent isolates each segment from the master file and feeds it into Seedance 2.0 along with the character reference for that shot. As one production put it: "I just uploaded the entire voiceover and told the agent that this is the part for this shot. So the agent trimmed it on its own, it fed it into Seedance with the character reference and it just generated the entire lip synced clip."
Step 4 — Lock voice consistency before generating dialogue shots. Instruct the agent to keep the voice the same across every shot before it starts rendering — one prompt, applied globally. For any B-roll narration where the character is off-screen, ask the agent to clone the voice from a locked dialogue shot and reuse it; invideo's voice cloning runs a 2-second audio-similarity check and auto-regenerates if vocal drift exceeds 1.5%.
Step 5 — Generate in batches of 5 and QC against the waveform. Run lip-sync clips in batches of 5 — small batches keep credits in check and let you catch a mismatched trim early. Review each clip against its source audio segment; if the mouth lags or the trim caught half a syllable, regenerate just that shot with a tightened in/out instruction ("start the segment at 'I tried', not at the breath before it"). Pixverse Lipsync handles the mouth animation pass inside the pipeline.
Step 6 — Stitch in your editor. Pull the locked lip-synced clips into Premiere Pro or invideo's Slate timeline in shot order. Because each segment came from one master VO, the timing across cuts already matches — no audio drift between shots.
Beyond auto-trim: across documented productions running this workflow, lip-synced UGC ads land around $70–$145 per ad (570 credits for a full localization, ~$125 average per English UGC ad), with ~85% clip rejection baked into those numbers — the per-shot trimming itself adds no extra credit cost because the agent reuses one VO file across every generation.
Watch some of these to see what works for you:
I just uploaded the entire voiceover and told the agent that this is the part for this shot. So the agent trimmed it on its own, it fed it into Seedance with the character reference and it just generated the entire lip synced clip.
— invideo's creative team, documenting the agent-directed voiceover trimming workflow