Best AI workflow for jewelry ad image generation with product consistency
Last updated August 1, 2026
For jewelry, build the base image in GPT-Image-2 to lock the aesthetic, then pass it through Nano Banana to lock the exact piece in — and run that pipeline inside the invideo agent so brand context, product references, and a CU/mid/wide consistency test govern every generation. One documented jewelry campaign held 100% product consistency this way at ~$2,400 across three ads.
Start by loading the invideo agent's context with everything the piece needs: brand guidelines, a product sheet showing the jewelry from every angle (front, side, back, macro), and a scale reference with a human hand holding it so the agent learns true size against skin. The invideo agent is an agentic creation tool that holds this context across every image and routes each generation to the right model — you don't switch platforms per shot.
Then run the two-model image pipeline. Build the base frame in GPT-Image-2 for aesthetic, composition, and lighting; then pass that frame through Nano Banana to lock the exact necklace, ring, or earring into it without intricate-detail drift. As Hridaye, invideo's creative director, puts it: "The best workflow is to build the base image with GPT image to get the aesthetic, then run Nano Banana to lock the exact necklace into it." Iterating on a single model when output looks generically AI is a dead end — when a shot drifts, render the same prompt across all available models in parallel and pick the winner, rather than burning credits on one.
Before committing to a full run, validate consistency at three focal distances. Generate one close-up, one mid, and one wide of the same piece and confirm the jewelry holds shape, stone count, and metal finish across all three. If it survives the three-distance test, the system is ready to scale; if it doesn't, fix the references first. Iteration is cheap here — invideo offers ~65% off image generation, so running 10 variations to find the right one is economically viable.
Lock each setup before moving on. Fully lock the first shot of a setup (lighting, angle, product placement) and every subsequent shot in that setup inherits the look. For multi-shot coverage of one beat — hero on velvet, hands clasping, gallery wall, close on the clasp — prompt a 3x3 coverage sheet of nine angles in one generation; three coverage passes can yield 20+ usable shots covering an entire ad setup. Use graphic matches (a wave curve echoing a necklace silhouette, water catching light like pavé) as the connective tissue that makes the set feel intentional rather than assembled.
The documented numbers from one jewelry campaign run end-to-end this way: brand film 160 images + 85 video clips generated, 13 used ($1,300); product film 85 images + 25 clips, 10 used ($625); anthem film 75 images + 37 clips, 12 used ($425) — $2,400 total for three ads holding 100% consistency on intricate jewelry. By around the fourth shot the agent had built enough taste context that notes like "more cinematic, low-angle workshop" landed in the first or second try.
A few things to enforce from the start: upload real reference images of any environmental props (workshop tools, displays, surfaces) — without them the model invents generic-looking versions; ask the agent for a sitrep ("what's locked, what's open — cast, location, lighting, product angle") before each new beat so you don't waste credits on unresolved decisions; and pull locked images into your editor as they finish, not at the end, so you can see the campaign cohere in real time.
Watch some of these to see what works for you:
The best workflow is to build the base image with GPT image to get the aesthetic, then run Nano Banana to lock the exact necklace into it.
— Hridaye, invideo's creative director