AI Ads

How does the two-stage GPT Image and Nano Banana pipeline work for AI product ads?

Last updated August 1, 2026

The two-stage pipeline generates each product-ad image twice: GPT-Image-2 builds the base image first — composition, environment, lighting mood, overall aesthetic — then Nano Banana runs a second pass that locks your exact product into that frame. Benchmarking showed GPT-Image-2 delivers the best aesthetic and Nano Banana the best product consistency, so each model handles the stage it wins.

invideo is an agentic video creation tool with all the current image and video models available, so both stages run inside one chat — the invideo agent passes outputs between models without you exporting anything. Here is the workflow in order:

1. Load product references before generating anything. Upload the product from every angle — front, side, back, detail close-ups, and every layer of packaging — plus a shot of the product held in a human hand so the models learn true scale. Feeding the invideo agent product images lets it understand the geometry and physics of the specific piece, which is what stage two locks against.

2. Stage one — GPT-Image-2 builds the base image. Prompt the shot's aesthetic: framing, environment, lighting mood, and any in-image text or design elements, where GPT-Image-2 also outperforms Nano Banana Pro. Don't fight this stage for perfect product detail — its only job is the overall look of the frame.

3. Stage two — Nano Banana locks the exact product into the frame. Pass the GPT-Image-2 base image plus your product references to Nano Banana and instruct it to place your exact product into the shot. In the documented jewelry production that established this pipeline, this is where the exact necklace was locked into the base image — Nano Banana held product detail that drifted on every other model tested.

4. Validate before committing a full run. Generate a close, a mid, and a wide with the product in all three; if the piece holds across the distances, the pipeline is ready to scale. If a shot still reads as too AI, don't keep iterating one model — ask the invideo agent to render the identical shot across the available models side by side and pick the winner. That six-model comparison is how this two-stage pairing was found in the first place.

5. Animate only locked frames. Once the two-stage image is approved, animate it with a video model — Seedance 2.0 or Kling, routed by the invideo agent — so video credits are spent only on frames where the product already holds.

Iteration economics make the two-pass structure practical: invideo runs a 65% discount on image generation, so producing 10+ image variations to get the right one costs little compared to video generation. The production that developed this workflow shipped three complete jewelry ads with 100% product consistency for roughly $2,400 total.

Watch some of these to see what works for you:

the invideo agent runs GPT Image then Nano Banana to lock jewelry product consistency

GPT image gave the best aesthetic, and Nano gave the best product consistency. The best workflow is to build the base image with GPT image to get the aesthetic, then run Nano Banana to lock the exact necklace into it.

— invideo's creative team

Share

More on AI Ads