How do you swap a product on a model in an AI image while keeping the face intact?
Last updated August 1, 2026
Swap the product in two discrete generation passes, never one: first remove the existing product while explicitly instructing the model to preserve all facial and skin texture detail, then add the correct product in a second pass using multi-angle product references. Documented result: a talent's necklace was removed and replaced in-agent with zero loss of face or skin fidelity.
Scope the edit to the product zone only — the core mistake is asking for the swap and the scene in one instruction, which forces a full-frame regeneration and drifts the face. invideo is an agentic creation tool with the current image models (GPT-Image-2, Nano Banana, Recraft) available, so the workflow below runs in one chat.
Step 1 — load multi-angle references of the new product. Upload the original talent image plus the replacement product from every angle: front, side, back, and a close-up; one production workflow used exactly 4 reference images per swapped item. For packaged goods, include a shot of the product held in a human hand so the model learns true scale. Multi-angle references are what let the swapped product hold 100% consistency instead of morphing into a generic version of itself.
Step 2 — remove the wrong product first, as its own pass. Instruct the invideo agent to delete the existing product while preserving all facial detail, skin texture, and pose. Keep the instruction narrow — nothing else in the frame changes. In one documented production, an existing talent image featured the wrong necklace; this removal pass took it out without touching face or skin fidelity.
Step 3 — add the correct product as a second pass. With a clean base, prompt the invideo agent to place the new product using your uploaded references. Splitting removal and addition means each generation solves one problem, so the model never has to reinvent the face while juggling the product.
Step 4 — route the product lock to the right model. GPT-Image-2 gives the strongest overall aesthetic, while Nano Banana delivers the best product consistency — so build or refine the base with GPT-Image-2, then run Nano Banana to lock the exact product into the frame. The invideo agent handles this routing, and every roster model is available inside it.
Step 5 — verify against the original. Compare face texture, skin tone, and proportions side by side with the source image before accepting. If an edit refuses to take, change the prompt's phrasing structure rather than repeating the same instruction — switching phrasing patterns is what forces specific physical edits in image generation.
If you need the same swap across a full video ad rather than a single image, the invideo agent also runs a reference-video product swap workflow — documented at about 115 credits (~$30) and 30 minutes per swap once the first one is set up.
Watch some of these to see what works for you:
The best workflow is to build the base image with GPT image to get the aesthetic, then run Nano Banana to lock the exact necklace into it.
— invideo's creative team