How many reference images do you need for consistent AI outfit and fabric generation?
Last updated August 1, 2026
Upload a minimum of 4 reference images per outfit — front, side, back, and a fabric close-up — and add a fifth showing the fabric worn on a person. Four locks garment construction; the close-up carries weave and texture; the on-model shot teaches drape. One documented production used exactly this 4-image set per outfit swap.
Build a 4-image set per outfit: a front shot, a side shot, a back shot, and a fabric close-up. A documented outfit-swap production uploaded exactly these 4 reference angles per new outfit to generate accurate swaps, and invideo's full fabric-consistency spec adds a fifth image — the fabric worn on a person — so the model sees how the material actually drapes and moves on a body.
Each image does a distinct job, which is why the count matters. Front, side, and back lock the garment's construction and silhouette so it doesn't morph between shots. The fabric close-up is what carries weave and color accuracy — evaluate every generated clip against weave accuracy, color accuracy, and how the garment interacts with body and environment. The worn-on-person shot sets drape and movement behavior that flat product shots can't communicate. Community fashion workflows report holding look consistency from as few as 3 reference images per garment, but 3 typically leaves fabric texture underspecified — the close-up is the image most workflows skip and most need.
More images only help if they agree. References with mismatched lighting, conflicting angles, or different garment versions degrade output rather than improve it, so shoot the set in one session under one light. If you do use multiple stylistic references, extract one attribute from each conversationally — the pose from one, the lighting from another, the framing from a third — instead of asking the model to blend them, which produces averaged mush.
For a full campaign, compile the reference sets into a lookbook — multiple angles, close-ups, product-only and on-model shots for every garment — and load it into the invideo agent's context once, so every generation in the project inherits it without re-prompting. Pair the images with a short written description of how each fabric physically feels (texture, weight, reflectivity); images define what the fabric looks like, the text governs how it behaves in motion. With that reference set locked, a two-person team produced two clothing ads with 100% fabric consistency in 8 hours for about $600, and reported that fabric behavior and texture were largely correct from the first generations — only framing needed iteration.
Watch some of these to see what works for you:
It's important to give the agent a close-up of the fabric, a front shot, a side shot, a back shot, and the fabric worn on a person.
— invideo's creative team