Blog

MAI Image 2.5: Microsoft AI's Image Generation and Editing Model (2026)

Last updated August 7, 2026

MAI Image 2.5: Microsoft AI's Image Generation and Editing Model (2026)

MAI Image is Microsoft AI's in-house image family, from MAI-Image-1 (Oct 2025) to MAI-Image-2.5-Pro preview (Jul 2026). The 2.5 flagship unifies text-to-image with precise localized editing — facial identity preservation, in-image text updates — and ranked #2 Arena editing per Microsoft. Token-metered pricing from $19.50/M (Flash) to $106/M (Pro) image-output tokens. Available as MAI Image 2.5 in the invideo agent.

Updated August 2026

MAI Image is Microsoft AI's in-house image model family — the line Microsoft built to power Bing Image Creator, PowerPoint, and OneDrive with its own generator instead of relying solely on OpenAI's image models. The current flagship, MAI-Image-2.5 (June 2, 2026), is a unified generation-and-editing model: one model that both creates images from text and performs precise, localized edits — swapping an object, updating in-image text, changing a background — while preserving faces and leaving the rest of the frame untouched. It appears in the invideo agent's picker as MAI Image 2.5.

What is MAI Image and how did the family evolve?

Microsoft AI shipped the first model, MAI-Image-1, on October 13, 2025, debuting in the top 10 for text-to-image on LMArena and rolling out as a user-selectable option inside Bing Image Creator. The cadence since has been fast:

Model Released What changed
MAI-Image-1 Oct 13, 2025 Debut; top-10 t2i on LMArena; photorealistic lighting, reflections, landscapes
MAI-Image-2 Mar 19, 2026 #3 t2i on Arena (per Microsoft); better in-image text; MAI Playground + select API customers
MAI-Image-2-Efficient Apr 14, 2026 ~41% cheaper, 22% faster — the "production workhorse" tier
MAI-Image-2.5 (+ 2.5-Flash) Jun 2, 2026 Unified generation and editing in one model; Flash is the budget tier
MAI-Image-2.5-Pro Jul 23, 2026 Highest-quality tier, public preview

Five releases in nine months, and only one changed what the model does rather than how fast or cheaply it does it: 2.5, which folded editing into generation. That release, not the leaderboard debut, is the family's real event.

Sources: official microsoft.ai posts for 2, 2-Efficient, 2.5, and 2.5-Pro.

What makes MAI-Image-2.5 different: generation plus surgical editing

Most image models generate; a smaller set edit; few do both well in one model. Per the official 2.5 announcement, the model handles object replacement, in-image text updates, motion-blur removal, and background changes "without collateral damage" — edits stay local instead of subtly repainting the whole frame. Two capabilities stand out:

  • Facial identity preservation across edits — change the outfit, background, or lighting and the person still looks like the same person.
  • Visual reasoning about scene structure, lighting, and scale — edits respect perspective and shadows rather than pasting objects in.

Microsoft is explicit about the boundary: MAI-Image-2.5 is unified for generation and editing, but "not designed for multimodal tasks beyond image generation/editing" — it is not a general vision-language model.

How does MAI Image benchmark?

Per Microsoft's own launch post — first-party numbers, so read them as vendor-reported — MAI-Image-2.5 ranked #2 on Arena for image editing (ahead of Nano Banana 2.1) and #3 for text-to-image at its June 2026 release, gaining +75 ELO overall versus MAI-Image-2 and +107 ELO on text rendering. The text-rendering jump is notable given in-image text was already a stated focus of MAI-Image-2. Independent leaderboard positions shift monthly, so check current Arena standings before treating these ranks as live.

What does MAI Image cost?

Microsoft prices the family per million tokens rather than per image — the token-metered model familiar from LLM APIs, applied to images (official rates, as of the June–July 2026 announcements):

Model Text input /1M Image input /1M Image output /1M
MAI-Image-2.5 $5.00 $8.00 $47.00
MAI-Image-2.5-Flash $1.75 $1.75 $19.50
MAI-Image-2.5-Pro $5.00 $8.00 $106.00

Practical read: Flash's image-output rate is under half of 2.5's, making it the volume tier; Pro at $106/M is priced for final-quality hero output. Actual per-image cost depends on resolution (more pixels = more tokens).

Strengths — and the weaknesses Microsoft flags itself

Strengths (official posts): photorealistic lighting and reflections from the MAI-Image-1 lineage; class-leading localized editing with identity preservation; strong in-image text after the 2.5 ELO gains; a real budget tier in Flash.

Weaknesses. Unusually for a lab, Microsoft's own documentation concedes the model "may reflect training-data biases" and can produce "plausible but inaccurate or misleading visual details," with human review advised in sensitive contexts. Beyond the official caveats: weights are closed, benchmark claims are first-party, and token-metered pricing makes per-image costs harder to predict than flat per-image rates.

Which jobs suit MAI Image 2.5?

  • Iterative product and marketing shots — generate a scene, then edit the label text or swap the background without re-rolling the whole image; a strong fit for product photography workflows where the product must stay pixel-identical across variants.
  • People-centric creative — identity preservation makes it the pick inside an AI image generator workflow when the same face must survive wardrobe, background, and lighting changes.
  • Text-bearing designs — the +107 text-rendering ELO gain (per Microsoft) plus text-editing support suits AI design generator tasks like posters and social frames where a headline has to be legible and revisable.

Prompting MAI Image 2.5: what works

  • Generate, then edit — don't re-prompt from scratch. The model's edge is localized editing, so treat the first generation as a draft and issue targeted edit instructions ("replace the sign text with 'SUMMER SALE', keep everything else").
  • Name what must not change. Identity and scene preservation respond well to explicit anchors: "same person, same pose, new background: rooftop at dusk."
  • Exploit scene reasoning. Describe lighting and scale relationships ("soft window light from the left, product at table height") — the model is documented to reason about them rather than ignore them.

MAI Image, answered

Can I use MAI Image commercially?

Commercial use runs through Microsoft's platform terms (Foundry/Azure). Note that in the MAI-Image-2 era, commercial API access required an application, so review the current Foundry terms for your use case before shipping paid work. Images you generate through platforms that license the model are governed by that platform's terms.

Is MAI Image open source?

No. All MAI-Image models are closed weights, available only through Microsoft's surfaces (Bing Image Creator, PowerPoint, OneDrive, MAI Playground) and API access via Microsoft Foundry.

What's the difference between MAI-Image-2.5, Flash, and Pro?

Same family, three quality/cost tiers: 2.5 is the balanced flagship ($47/M image-output tokens), Flash is the fast budget tier ($19.50/M), and Pro (public preview since July 23, 2026) is the top-quality tier at $106/M.

Can MAI-Image-2.5 edit existing images, not just generate?

Yes — that's its defining feature. It performs localized edits (object replacement, text updates, background changes, blur removal) on existing images while preserving facial identity and untouched regions, per the official June 2026 announcement.

Where does Microsoft use MAI Image in its own products?

In production behind Bing Image Creator (generation), PowerPoint (image generation), and OneDrive (photo editing), per Microsoft's announcements — meaning the model is battle-tested at consumer scale.

Is MAI-Image-2.5 a multimodal model like GPT-image?

No. Microsoft states it is purpose-built for image generation and editing only, "not designed for multimodal tasks beyond image generation/editing."

Where can you try MAI Image 2.5?

You don't need a Foundry account or token math to use it. MAI Image 2.5 is part of the 200+ model lineup inside invideo — pick it in the agent when a job calls for edit-heavy, identity-safe images, and let the platform handle the metering. It sits in the picker next to every other major image family (browse the AI models index for the full set), and pairs naturally with the product photography and AI image generator workflows where its editing strengths do the most work.


Version history: first published August 2026, covering MAI-Image-1 (Oct 2025) through MAI-Image-2.5-Pro public preview (Jul 2026).

Share