
GPT Image 2 (April 2026) is OpenAI's flagship image model and the #1-ranked text-to-image model on both major arenas as of August 2026. It reasons before rendering — researching, planning layout, self-checking — which makes it the strongest model for text-heavy and structured images. Older family members are being retired: gpt-image-1 sunsets October 23, 2026, gpt-image-1.5 on December 1, 2026. The family runs inside invideo's agent.
Updated August 2026
GPT Image 2 is OpenAI's flagship image generation model, released April 21–22, 2026, and — as of August 2026 — ranked #1 on both major text-to-image leaderboards. It caps the GPT Image family, the line that replaced DALL-E: GPT Image 1 (March 2025), 1 Mini (October 2025), 1.5 (December 2025), then GPT Image 2. Its defining trait is that it reasons before it renders — researching the subject, planning the layout, and checking its own work — the first image model built this way. The family also has an expiry schedule most 2026 guides omit: half of it shuts down within months.
How did the GPT Image family evolve?
GPT Image 1 (March 25, 2025) launched inside ChatGPT as "4o image generation" and became a cultural event: 130 million+ users created over 700 million images in the first week, driven largely by the viral Studio Ghibli-style portrait trend (Wikipedia: GPT Image). OpenAI's first autoregressive image model — a break from DALL-E's diffusion approach — it excelled at instruction-following and in-image text, but had a warm yellow cast, premature cropping, and weak multi-face scenes. The API version followed on April 23, 2025.
GPT Image 1 Mini (October 6, 2025), announced at DevDay 2025, is the cost tier — roughly 80% cheaper than GPT Image 1, with the same three output sizes but one hidden trade-off covered below.
GPT Image 1.5 (December 16, 2025) shipped in the API and rolled out globally in ChatGPT as the relaunched "ChatGPT Images." It generated up to 4× faster, fixed the cropping and yellow tint, added face and logo preservation via the input_fidelity parameter, and cost about 20% less than GPT Image 1, with noted regressions in some art styles (Wikipedia).
GPT Image 2 (April 21–22, 2026) is a new architecture, not a 4o derivative. After anonymous testing on LMArena in early April as maskingtape-, gaffertape-, and packingtape-alpha, it launched in the API and Codex on April 21 and reached all ChatGPT users as "ChatGPT Images 2.0" a day later (OpenAI Developer Community, April 2026).
How do GPT Image 2's specs compare to 1.5 and earlier?
| Spec | GPT Image 1 / Mini / 1.5 | GPT Image 2 |
|---|---|---|
| Output sizes | Fixed: 1024×1024, 1536×1024, 1024×1536 | Arbitrary, per Azure: 16 px increments, long edge ≤3,840 px, ratio ≤3:1 (see note) |
| Quality tiers | low / medium / high (Mini defaults medium) | low / medium / high; low latency-optimized |
| Images per request | 1–10 | 1–10 via API; "up to 8 coherent images per prompt" |
| Editing | Inpainting with PNG mask, multi-image input, variations | Same, plus multi-image batch editing |
| Face/logo preservation | input_fidelity on 1 and 1.5 — not on Mini |
Supported, advanced face preservation |
| In-image text | Latin reliable; CJK/Arabic weak on GPT Image 1 | Strong multilingual rendering incl. Japanese, Korean, Chinese, Hindi, Bengali |
| Output formats | PNG (default, transparency) or JPEG; streaming partial previews; typically 10–30 s | Same |
Sources: Azure AI Foundry image docs, OpenAI Developer Community, August 2026.
A resolution discrepancy worth knowing: OpenAI's launch copy says GPT Image 2 generates "up to 2K," while Azure's current documentation specifies arbitrary resolutions up to a 3,840 px long edge — 4K-class. Likely a post-launch expansion documented on Azure first, but as of August 2026 the two official sources genuinely disagree; test anything above 2K before relying on it.
Which GPT Image model should you actually use?
| Picker entry | What it is | Pick it when |
|---|---|---|
| GPT Image 2 | Flagship: agentic reasoning, best text rendering, multilingual, batch editing | Almost everything — finals, layouts, text-heavy work |
| GPT Image 1.5 | Fast, fixed 1.0's flaws, input_fidelity support |
Legacy pipelines only — shuts down December 1, 2026 |
| GPT Image 1 | The original, now outclassed | Nothing new — shuts down October 23, 2026 |
| GPT Image 1 Mini | ~80% cheaper; medium quality default | Bulk drafts and prototyping where cost beats polish |
| Heron Alpha | In invideo's picker as a lighter GPT-Image-2-family option | Experimental — see note below |
| Hawk Alpha | In invideo's picker as a lighter GPT-Image-2-family option | Experimental — see note below |
About Heron Alpha and Hawk Alpha: both appear in invideo's roster as lighter GPT-Image-2-family candidates, but OpenAI has published no documentation for either codename as of August 2026 — the only confirmed pre-release codenames were LMArena's "tape" trio, and no documented mini variant of GPT Image 2 exists. Treat them as in-product experimental options, not documented OpenAI releases.
The Mini trade-off most comparisons miss: GPT Image 1 Mini does not support input_fidelity. Downshift to Mini on an editing workflow and face/logo preservation silently degrades — no error, just worse likenesses (Azure docs).
What makes GPT Image 2 different from other image models?
It reasons first. GPT Image 2 applies O-series-style agentic reasoning to image generation: given "an infographic explaining how espresso extraction works," it researches the subject, plans the composition, renders, and self-checks before returning a result. That is why its diagrams tend to be correct-by-construction — labels in the right places, clock faces showing the stated time, readable speech bubbles in comic panels — rather than plausible-looking approximations (OpenAI Developer Community, April 2026).
The blind-test record backs the approach. On the July 10, 2026 Arena.ai leaderboard, gpt-image-2 sits #1 at 1385 Elo — roughly 83 points clear of #2 — and Artificial Analysis also ranks it #1 at 1339 Elo (August 2026; different Elo scales, same verdict). At launch, OpenAI claimed a "+242-point lead in Text-to-Image" over the previous leader.
Where does GPT Image 2 fall short?
Reasoning costs time — a poor fit for real-time generation, where the low tier or Mini serve better. The knowledge cutoff is December 2025, so recent subjects lean on web-search integration. And on multi-reference character consistency, Google's Nano Banana Pro still leads: the Gemini-native line accepts up to 14 reference images, which GPT Image 2 does not match as of August 2026.
Which GPT Image models are shutting down?
The deprecation cascade, dated (as of August 2026): DALL-E 3 was retired on Azure on March 4, 2026, and DALL-E 2/3 were removed from the OpenAI API on May 12, 2026 (Azure docs).
gpt-image-1is scheduled for API removal on October 23, 2026, andgpt-image-1.5on December 1, 2026 — an unusually short ~11-month flagship lifespan for 1.5. The migration target for everything is GPT Image 2.
Verify the exact dates against OpenAI's deprecations page before committing a production pipeline; schedules occasionally shift.
How much does GPT Image cost?
OpenAI prices image generation in tokens. Official API rates, as of August 2026:
| Model | Text input /1M tokens | Image input /1M tokens | Image output /1M tokens |
|---|---|---|---|
| gpt-image-2 | $5 | $8 ($2 cached) | $30 |
| gpt-image-1.5 | $5 | $8 | $32 |
| gpt-image-1 | $5 | $10 | $40 |
| gpt-image-1-mini | $2 | $2.50 | $8 |
Source: OpenAI Developer Community announcement, August 2026. Per-image cost scales with size and quality — roughly half a cent (small/low) to about $0.21 (large/high) on GPT Image 2. A detail few noticed: GPT Image 2 dropped the text-output token charge entirely and cut image output about 6% versus 1.5 ($32 → $30 per million tokens).
How do you prompt GPT Image 2?
- State intent and format, not composition. The reasoning step handles layout. "A four-panel comic about deadline panic" or "an onboarding infographic for a budgeting app" outperforms long positional instructions.
- Spell out exact strings. Text that must appear verbatim — headlines, UI labels, multilingual captions — goes in quotes in the prompt. GPT Image 2 renders it, including CJK and Hindi scripts.
- Use
input_fidelity: highfor faces and logos. On edits where a person or brand mark must survive, this parameter is the difference between preservation and reinterpretation — and it is unavailable on Mini. - Preview cheap, finalize expensive.
quality: lowwith streaming partials for iteration;highonly for finals, since pricing scales with output tokens.
What can you make with GPT Image 2?
- Infographics and diagrams. The reasoning-driven accuracy is the family's signature skill — a natural fit for infographic work where labels and data callouts must be right.
- Thumbnails. Reliable in-image text plus strong faces suit high-CTR thumbnail production; draft in bulk on Mini, rerender winners on GPT Image 2.
- Posters and text-heavy layouts. Multilingual typography and layout planning cover poster formats, from gig posters to localized campaign art.
GPT Image FAQ
Can I use GPT Image 2 images commercially?
Yes. OpenAI's Terms of Use assign output ownership to the user, and commercial use is permitted. Review the current terms before relying on this for client work.
Do GPT Image images carry a watermark or C2PA metadata?
No visible watermark. OpenAI has attached C2PA provenance metadata as policy since DALL-E 3, and this reportedly carries into the GPT Image line — but the metadata can be stripped by re-encoding, so it is provenance, not protection. Check OpenAI's current help documentation for exact scope; details are not fully confirmed as of August 2026.
What sizes can GPT Image 2 generate?
GPT Image 1, Mini, and 1.5 output three fixed sizes: 1024×1024, 1536×1024, and 1024×1536. GPT Image 2 supports flexible resolutions — per Azure's docs, dimensions in 16 px increments up to a 3,840 px long edge and 3:1 ratio, though OpenAI's launch copy says "up to 2K." Test above 2K before depending on it.
Which GPT Image variant should I use for faces?
GPT Image 2 with input_fidelity: high, or GPT Image 1.5 while it lasts. Never Mini — it lacks input_fidelity entirely, so face preservation degrades without warning.
Is DALL-E still available?
No. DALL-E 3 left Azure on March 4, 2026, and DALL-E 2/3 left the OpenAI API on May 12, 2026. The GPT Image family is OpenAI's only image line as of August 2026.
When do GPT Image 1 and 1.5 shut down?
Per the schedule as of August 2026: gpt-image-1 on October 23, 2026, gpt-image-1.5 on December 1, 2026. GPT Image 2 is the migration target for both.
How do you run GPT Image without an API key?
No OpenAI API key is needed to run this family. GPT Image 2, 1.5, 1, and 1 Mini all sit in the image roster of invideo's agent — alongside 200+ other models including Nano Banana Pro, Veo 3.1, and Sora 2 — which can route text-heavy layouts to GPT Image 2, hand other shots to whichever model suits them, and carry the stills into video. The GPT Image 2 hub on invideo covers the model in the platform's context, and the AI image generator workflow is the shortest path from prompt to finished image.
Version history: GPT Image 1 Mar 25, 2025 (API Apr 23) → 1 Mini Oct 6, 2025 → 1.5 Dec 16, 2025 → GPT Image 2 Apr 21–22, 2026. DALL-E retired Mar–May 2026; gpt-image-1 sunsets Oct 23, 2026; gpt-image-1.5 Dec 1, 2026.