Why should you upload a reference video when briefing an AI agent on a complex ad hook?
Last updated August 1, 2026
Upload a reference video when a hook's timing, cut rhythm, snap-to-cut energy, and staged camera move can't be reliably described in text. The video locks the exact beat structure, pacing, and visual cadence the invideo agent needs — eliminating the guesswork that text-only briefs leave behind on complex hooks.
Text breaks down on complex hooks because hooks are precisely the elements text encodes worst: the millisecond a hand enters frame, the match-cut from messy to spotless on a snap, the energy shift between the pre-reveal and the product moment. invideo is an agentic video creation tool, and the invideo agent reads uploaded reference videos to decode this structure directly — pulling cut points, per-scene framing, and pacing into its working context before any generation begins.
What a reference video communicates that text cannot
Hook timing and the energy shift. The 1.5-second scroll-stopping window, the exact frame the snap lands, the cut from before-state to after-state — these live in the video's frame timing, not in adjectives. Hridaye, invideo's creative director, flagged this directly while briefing a match-cut hook: "I pause here because I'm trying to understand if the agent is being able to visualize the hook that I've written in text or does it need a reference video." When the answer is yes, the reference goes in.
Cut rhythm and shot count. When given a reference, the invideo agent detects every cut automatically — in one documented production it identified 9 cuts in the reference ad and extracted one representative frame per scene, then preserved 7 shot beats across localized versions. Text would have required listing every cut by hand and still missed the cadence.
Camera language and staging. The agent extracts and stores production rules from the upload — cut rate, camera style, energy, ending pattern — and organizes them under structured headings before generating anything. You're effectively handing it a director's playback of the hook instead of asking it to imagine one.
Voiceover pacing and music structure. The agent will transcribe the reference's voiceover (even across languages), deconstruct the music bed, and adapt both — locking the rhythm the hook depends on rather than approximating it.
How to brief the upload so the agent uses it correctly
Upload the reference and ask the invideo agent to break it down into named beats with timestamps, locations, function, and emotional logic before any generation. This forces structure onto what the video is actually doing.
Get the agent's planned execution back as text — its read on cut count, hook duration, product-reveal moment, CTA cadence — and confirm it before spending video credits.
Lock the shot breakdown (shot number, duration, description, super text) as the contract for production. Generation only starts after this is signed off.
If the reference has burned-in captions, do NOT attach the video itself at generation time — pass the agent the extracted character sheet, location sheet, and shot-flow instead. Otherwise the model will reproduce the old captions in the new shots.
Your reference video should cover
The full hook beat (the gesture, expression, or visual tension that earns the scroll-stop)
The cut that resolves it (match-cut, location change, product reveal)
The pacing signal between hook and body (where the energy modulates)
The product moment and CTA cadence (so timing carries through to the end)
If any of those four elements live only in your head, a text brief will guess at them. The upload turns guesses into constraints.
Watch some of these to see what works for you:
I pause here because I'm trying to understand if the agent is being able to visualize the hook that I've written in text or does it need a reference video.
— Hridaye, invideo's creative director