How do you communicate accurate product scale to AI for video ad generation?
Last updated August 1, 2026
Communicate scale to AI by giving it visual anchors, not adjectives: upload reference images that include a human hand holding the product, lock a separate keyframe per product size when sizes vary in one ad, validate the product at close/mid/wide before any batch run, and keep a recognizable scale cue (hand, palm, table edge) visible in the frame.
Scale drift happens because video models have no innate sense of real-world dimensions — they infer size from whatever cues sit in the frame. Fix it at the asset stage, not in prompts.
invideo's agent is an agentic video tool that holds project context across every generation, so any scale references you load once apply to every shot the invideo agent renders afterward.
Build a product sheet with a hand-held scale reference. Before generating anything, upload every product angle — close-up, front, side, back — plus a shot of a human hand holding the product and its packaging. One documented production framed it directly: "I give the agent the reference images of the hand holding the sachet, a few gummies in the palm, so it understands the actual size of all of the layers of your packaging compared to the human hand." This is how you solve product consistency at the source, before a single video credit is spent.
Validate scale at three focal distances before batching. Generate one close-up, one mid, and one wide shot with the product visible in each, and confirm the product holds the same proportions across all three. If the piece drifts between distances, fix the reference set — don't commit a full generation run on a scale that isn't holding.
Use a separate locked keyframe per product size when sizes vary in one ad. When a character interacts with products of significantly different sizes in the same ad (a small sachet in one shot, a larger pouch in another), build and lock a distinct keyframe per product scale rather than one unified keyframe — this prevents the model from confusing scale relationships during generation and cuts re-gens.
Generate the still first, then animate from the locked frame. Iterate cheaply on a static image until the scale relationship between product, hand, and environment reads correctly, then spend video credits only on animating that locked frame. One creative director put the discipline plainly: "I only spent video credits on locked frames." Image generation runs in seconds on the invideo agent, so iteration to a correct-scale keyframe costs almost nothing.
Keep the scene compositionally simple around the product. Fewer competing scale cues means less ambiguity for the model — one hero object, a clean background, and one recognizable scale anchor (hand, palm, table edge, retail shelf) in frame. Save dense environmental composition for shots where the product's scale is already established.
Route to the right model for product-faithful rendering. For the hero scale-establishing still, Nano Banana locks exact product geometry into a base image, and GPT-Image-2 holds aesthetic and text/packaging detail; for the motion clip, Seedance 2.0 reference-to-video carries the product and scale context from your locked keyframe into the animation. The invideo agent has all of these available and routes per shot, so you're not picking a platform per model.
Add a sitrep check before locking. Ask the invideo agent for a sitrep — what's locked, what's open on product, scale reference, and framing — before approving the batch. It surfaces any scale decision that hasn't actually been resolved and prevents wasted regenerations downstream.
These techniques compound: one production run holding 100% product consistency across a full jewelry campaign came in at ~$2,400 for three ads, with the scale-and-consistency work front-loaded into the product sheet and the multi-distance validation. Get the references right once and every subsequent ad inherits the same scale logic.
Watch some of these to see what works for you:
I give the agent the reference images of the hand holding the sachet, a few gummies in the palm, so it understands the actual size of all of the layers of your packaging compared to the human hand.
— invideo's creative team, on the Product Sheet method