AI Ads

What reference images should I upload for AI fabric ad generation?

Last updated August 1, 2026

Upload five image types per garment: a fabric close-up showing the weave, a front-on flat or on-model shot, a side shot, a back shot, and the garment worn on a person. Add a hand-holding-product reference for scale, and colorway variants if multiple colorways exist. Garment photos beat fabric swatches every time.

Start with the five-image core set the invideo agent needs to lock fabric behavior across every generated shot:

Fabric close-up — a tight macro of the weave so the model can resolve thread structure, sheen, and surface texture. Without this, fabric reads as plastic or generic AI cloth.

Front shot — full garment on a neutral background, wrinkle-free, evenly lit. This is the hero silhouette the agent will return to for every wide and mid shot.

Side shot — establishes drape and how the fabric falls along the body's profile. Critical for movement shots.

Back shot — locks the garment's full 3D geometry so the agent doesn't invent the rear when characters turn.

Worn-on-person shot — shows how the fabric sits on a real body, with natural folds, tension points, and the way it moves with the wearer. This is what teaches the agent the difference between a stiff render and a garment that actually behaves like the material.

For products with packaging (gummies, sachets, small goods), add a hand-holding-product reference — palm + product + packaging layers in one frame — so the agent knows true scale relative to a human hand and doesn't drift on size across shots. In one documented production, this single reference image was the difference between repeated scale errors and clean first-try generations.

Where multiple colorways exist, upload one set per colorway and name the files clearly — the invideo agent reads filenames and routes each variant to the correct shot.

Pair every image with a one-line descriptor when you upload it: "close-up of 100% linen weave, natural drape", "front, side, back of cropped jacket — heavy cotton twill", "hand holding 28-day pouch for scale". The agent uses the words to interpret what's IN the image — without that, it may pull pose or framing when you wanted texture.

Image specs that matter: high-resolution stills (sharper input = sharper output), neutral or clean background, even lighting with no heavy filters, color-accurate. Garment photos consistently outperform fabric swatches — swatches give the agent material but no information about how it falls, folds, or interacts with a body.

If you also have a reference film or winning ad, upload it separately for camera language and pacing — not as a fabric reference. Mixing the two confuses what the agent should extract from each. Tag what you want from each reference in the chat: "pull pose from image 1, lighting from image 2, framing from the film reference."

Once the set is uploaded to the invideo agent's context tab, it persists across every shot in the project — you upload once, and fabric stays consistent across sun, wind, touch, and motion without re-prompting. In one production, two complete clothing ads with 100% fabric consistency were produced in 8 hours by 2 people, with the creator never re-prompting fabric behavior shot-by-shot because the reference set lived in context.

Watch some of these to see what works for you:

Full tutorial: how reference images lock fabric consistency in AI fashion ads
See how garment photo uploads translate into consistent AI fashion ad output
How to pass multiple garment references to an AI agent without blending them

It's important to give the agent a close-up of the fabric, a front shot, a side shot, a back shot, and the fabric worn on a person.

— invideo's creative team

Share

More on AI Ads