How do you use a coverage sheet or 3x3 grid to generate multiple camera angles in one AI video pass?
Last updated August 1, 2026
A coverage sheet is a 3x3 grid of nine shots of the same subject — varied by angle, framing, and lens — generated as one image in a single pass. You prompt the invideo agent for the grid, lock the cells that work, then animate each cell as its own clip. One documented fashion production got 20+ usable shots from three grid generations.
Run it as a five-step sequence inside one invideo agent project so the grid inherits your character sheet, wardrobe lock, and visual rules automatically.
1. Lock the subject before you ask for the grid. Generate and lock the character sheet (front/side/back) and the keyframe location first. The grid only works if the agent already knows who and where — otherwise each of the nine cells drifts. invideo's creative director Hridaye describes the underlying principle directly: "One clean image works as the anchor for every future gen."
2. Define your nine cells as one narrative beat, three tiers deep. Pick a single moment from your shot list (e.g. "model walking into frame, linen dress catching wind") and break coverage into three tiers of three: wide/establishing (full body, environment), medium/core (waist-up, hero action), tight/detail (fabric macro, hands, face). Write the nine cells out explicitly — angle + framing + lens behavior per cell — before prompting. Decomposed reference pulling helps here: tell the agent which attribute to pull from where ("pose from cell 4, lighting from the keyframe, framing from the reference still").
3. Prompt the grid as one image, in your film's aspect ratio. Ask the agent for a 3x3 contact sheet of the locked subject, with each cell labeled by its angle and framing, all sharing the same character, wardrobe, lighting direction, and color grade. Route it to a strong reference-aware image model — GPT-Image-2 holds character and location reference cleanly, Nano Banana locks exact product detail across cells, Recraft holds skin texture for casting-tight shots. The invideo agent picks the right one for the cell type; you don't have to swap platforms.
4. Curate, then audit for editorial gaps. Review the nine cells, mark which hold and which drift. Before you animate, query the agent for missing shot types — back shots, product still lifes, cropped fragments, seated/reclined poses, empty atmospherics, duo compositions. A standard lookbook brief omits these; the editorial gap audit surfaces them and you regenerate a second grid to fill them. Three grid passes typically yields the 20+ usable shots a full beat needs.
5. Animate each locked cell as its own clip. Feed each cell into Seedance 2.0 as a reference frame with explicit camera-movement instructions per cell (push-in on the detail cells, lock-off on the wides, slow arc on the mediums). Generate in batches of five clips to keep credit spend tight and iterate fast. If Seedance's 15-second cap forces splits, the invideo agent decomposes the batch automatically. Assemble in Slate inside invideo or pull to Premiere Pro as clips lock.
Why the grid pays off: image generation runs at heavy discount versus video, so a nine-cell grid costs a fraction of generating nine separate video clips and discovering drift in the edit. One documented fashion campaign ran 40 editorial stills and 30 motion clips end-to-end for ~$150 in 3–4 hours using exactly this pattern — grid first, animate the locked cells.
One probe before you scale. Generate one full-resolution cell (character + garment + set + skin + pose all present) before committing the rest of the grid to a batch. If that single shot holds, the system is ready; if it drifts, fix the lock before you burn credits on nine cells that all carry the same flaw.
Watch some of these to see what works for you:
Three generations, 20 plus usable shots, the whole ad covered.
— Hridaye, invideo's creative director