AI Ads

How do you use a reference ad to generate a new AI video ad in the same style?

Last updated August 1, 2026

Upload the reference ad to the invideo agent and have it deconstruct the ad before generating anything: it transcribes the script, detects every cut (9 cuts auto-detected in one documented ad), extracts one frame per scene, and writes production rules — camera style, pacing, cut rate, ending pattern — for the new ad. Review those rules, lock a new shot breakdown, then generate.

Upload the reference ad as a video attachment and instruct the invideo agent to break it down before any generation. invideo is an agentic video creation tool with all the current models available, so the analysis, image generation, video generation, and assembly all happen in one project. The invideo agent transcribes the script and identifies the structural elements worth keeping, detects the exact timestamp of every cut, and extracts one representative frame per scene — in one documented production it detected 9 cuts automatically and used one frame per scene to guide recreation. If the reference is in a foreign language, ask it to translate the voiceover and deconstruct the music structure so both can be adapted for the new ad.

Next, review the style rules the invideo agent extracts. It organizes them under structured production headings — Visual Standard, Pacing & Camera, cut rate, camera style, on-screen text, energy, and ending pattern — and saves them in the project's context tab, so every subsequent generation follows them without re-prompting.

Then lock the new shot breakdown. The invideo agent returns a shot-by-shot table derived from the reference — shot number, duration, shot description, and super text — and you review and approve it before any image or video generation begins. Brief it the way you would a crew: state what stays identical (edit structure, music bed, pacing, shot beats) and what changes (product, character, market, on-screen text). In one documented localization, 7 shot beats from the original were preserved exactly while only cast, location, voiceover, and UI text changed.

Generate still images before video. Iterate each frame cheaply until it matches the reference's composition, lock it, and spend video credits only on locked frames — this is the primary cost-control step. Seedance 2.0 animates the locked frames into clips, and because every roster model (Kling, Veo, Seedance 2.0) runs inside invideo, the invideo agent routes each shot to whichever model holds character and product consistency best for that shot.

One warning during clip generation: if the reference ad has burned-in captions, do not attach the reference video to the generation prompts — the model will reproduce the old captions in your new footage. Direct generation using the extracted frames, the locked shot breakdown, and your reference sheets instead.

Finally, assemble in the same edit order as the source ad. Pick the best generated version of each segment, lock it, and stitch — either in Slate, invideo's built-in timeline editor, or in Premiere Pro — keeping the original sequence so the pacing that made the reference work carries over.

Once the first recreation is locked, the workflow repeats at much lower cost: a product swap on the same reference structure ran 115 credits ($30) and took about 30 minutes after the first, a full character-and-language localization ran 570 credits ($145), and combined output across both approaches reached 8–10 variations of a single winning ad in one day with one senior creative running the system.

Watch some of these to see what works for you:

Full tutorial: scale a winning ad into new characters, languages, and markets
Localize UGC ads end-to-end — including how to avoid copying reference captions

The agent watches the video and comes back with a plan. It tells me what stays the same and what needs to change.

— invideo's creative team

Share

More on AI Ads