When AI jewelry looks plastic, the fix is not more prompting on one model — it's a multi-model pipeline plus a product sheet anchor. Render the same shot across every available video and image model, build the base aesthetic in GPT-Image-2 and lock the exact piece in with Nano Banana, and validate at close, mid, and wide before committing the run.
Start by diagnosing which failure you're actually seeing — plastic metal sheen, dead gemstones with no refraction, melted prongs and engraving, inconsistent shadows, or color drift between shots. Each one has a specific fix, and stacking them is what gets jewelry to read as real.
The invideo agent is an agentic video tool with every current image and video model — GPT-Image-2, Nano Banana, Recraft, Runway, Veo, Kling, Seedance 2.0 — available inside one project, so the moves below run without switching tools.
Render the same shot across multiple models and pick the winner. Iterating harder on one model that's producing plastic output is a dead end. Ask the invideo agent to render the identical prompt across each available image or video model simultaneously, then choose the result that holds the metal and stones. In one documented jewelry production, six models were compared on the same close-up — Kling V3 Omni, Wan 2.7, Seedream 5.0 Lite among them — and the selected output was visibly less "AI" than any single-model iteration loop produced.
Build the base in GPT-Image-2, then lock the exact piece in with Nano Banana. This dual-model pass is the workflow that solved a documented three-ad jewelry campaign with intricate pieces. GPT-Image-2 establishes the aesthetic — lighting, environment, composition. Nano Banana then takes that base image plus reference photos of the actual product and locks the real piece into it, preventing the prong count, stone cut, and setting from drifting shot to shot. Hridaye, invideo's creative director, put it directly: "The best workflow is to build the base image with GPT image to get the aesthetic, then run Nano Banana to lock the exact necklace into it."
Build a product sheet before generating anything. Upload every angle of the piece — top, side, three-quarter, back, macro of the setting — plus one reference photo of a hand holding the piece so the model learns true scale. Without this, AI invents generic-looking jewelry and gets the size relative to skin, hand, or neck wrong. This is what fixes scale drift and contact-point errors (jewelry hovering off skin, clasps that don't close) at the source rather than in post.
Validate at three focal distances before the full run. Generate one close-up, one mid, and one wide of the piece in context. If the stones refract, the metal carries micro-reflections, and the engraving holds at every distance, the system is ready to scale. If close-ups soften or color drifts wide, fix it now — propagating a broken setup across 20 shots wastes credits.
Lock the first shot of each setup completely. When you're shooting a setup (workshop, gallery, on-model), fully nail the first shot — lighting, reflections, depth — before generating any others in that setup. Subsequent shots inherit the look. Pair this with graphic matches across cuts (a wave curve echoing a necklace silhouette, water catching light like pavé) so the film reads intentional rather than assembled.
Upload real brand reference, not generic descriptions. If the piece lives in a workshop or boutique context, give the agent real reference photos of that environment. Without them, the AI invents generic-looking tools and surfaces; with them, the environment matches brand reality. Same principle for the piece itself — actual product photography in context teaches the model what "correct" looks like.
One documented jewelry campaign ran three full ads (brand film, product film, anthem film) using exactly this stack — multi-model rendering, GPT-Image-2 plus Nano Banana, product sheet anchoring, three-distance validation — for approximately $2,400 with 100% product consistency across all three films. By shot four, the agent had absorbed enough taste context that creative directions landed first or second try.
Beyond the pipeline itself: AI image models still lack true optical simulation of caustics, sub-surface scattering through diamonds, and physically accurate metal BRDF. The ceiling fix is hybrid — AI base plus a manual inpaint or composite pass on the hero reflections and facets in a finishing tool. The moves above get you to the point where that last pass is touch-up, not rescue.
Watch some of these to see what works for you:
Jewelry is the hardest category to pull off with AI. The products are super intricate and using raw models means your product changes shot to shot. But setting up and using an AI agent with proper direction solves this issue.
— Hridaye, invideo's creative director