Microsoft's MAI-Image-2.5 Hits No. 2 on Leaderboard With Precise Localized Editing
Microsoft's MAI-Image-2.5 ranks No. 2 for image editing on Arena, with a unified generation and editing API now live on Foundry and OpenRouter

- MAI-Image-2.5 is Microsoft's new flagship image model, ranking No. 2 for image editing and No. 3 for text-to-image on Arena's human-preference leaderboard.
- It delivers +75 Arena points over MAI-Image-2, with the biggest gains in text rendering (+107) and cartoon/anime/fantasy (+90).
- Key capability: localized editing that changes one object without touching the rest of the image, plus face and identity consistency across edits.
- Ships in two variants: MAI-Image-2.5 (max fidelity) and MAI-Image-2.5-Flash (faster, ~41% cheaper on image output at $19.50/1M tokens).
- Available now on the MAI Playground, Azure AI Foundry, and OpenRouter; also live inside PowerPoint and rolling out to OneDrive.
- Proprietary model (not open weights) from Microsoft's in-house MAI team, part of a seven-model family launched at Build 2026.
Microsoft just made a serious move in the image generation race. MAI-Image-2.5, the latest model from the company's in-house MAI team, launched as part of a seven-model family announced at Build 2026. It ships in two variants, targets production image workflows, and immediately staked out a top-three position on the most-watched human-preference leaderboard in the space.
Not just another image model
MAI-Image-2.5 is Microsoft AI's updated flagship image-generation model, purpose-built for high-quality text-to-image generation and precise, controllable image-to-image editing at production scale. What makes this release different from the usual "better photorealism" announcement is the editing story. What makes MAI-Image-2.5 interesting is not just generation quality but editing precision: it supports localized edits that change one object without disturbing the rest of the image, and it preserves facial identity across pose and expression changes.
Most image models regenerate the entire frame when you ask for a small change, which means faces shift, backgrounds drift, and products lose their exact look. Localized editing that leaves the rest of the image untouched is what makes a model actually usable for e-commerce catalogs and brand assets, where consistency is the whole game.
Where it lands on the leaderboard
MAI-Image-2.5 now ranks No. 2 on Arena's Image Edit leaderboard, ahead of Nano Banana 2.1. Arena (arena.ai) is a blind human-preference leaderboard where real users vote on head-to-head image comparisons without knowing which model produced which output. It's the closest thing the field has to an unbiased quality signal.
MAI-Image-2.5 delivers an overall +75 point improvement over MAI-Image-2, with the largest gains in Text Rendering (+107) and Cartoon, Anime & Fantasy (+90). The text rendering jump is particularly notable. Text rendering has historically been a weak spot for image models, so a large jump there is directly useful for anything involving product labels, signage, or UI mockups.
Four things it's actually good at
- Text-to-image quality: MAI-Image-2.5 produces more detailed, coherent images from prompts, with stronger text rendering, product imagery, and prompt adherence.
- Scene understanding: The model understands scene structure, lighting, scale, and spatial relationships, helping it make edits that fit the image context, such as adding an object with the right perspective and shadows.
- Localized editing: MAI-Image-2.5 supports precise, localized edits, from replacing an object or updating text to removing motion blur, without changing the rest of the image.
- Identity consistency: It introduces identity and character consistency across stylization, pose, and layout, and structured document and diagram generation that produces PowerPoint-ready visuals and slides.
Two variants, one API
Microsoft launched two versions at once: MAI-Image-2.5 for maximum fidelity, and MAI-Image-2.5-Flash for fast, scalable production workloads where speed and cost matter more than the last few points of quality. The pricing is token-based, split across text input, image input, and image output.
| Price (per 1M tokens) | MAI-Image-2.5 | MAI-Image-2.5-Flash |
|---|---|---|
| Text input | $5.00 | $1.75 |
| Image input | $8.00 | $1.75 |
| Image output | $47.00 | $19.50 |
Image output is the dominant cost for generation workloads. Flash brings that from $47 down to $19.50 per million tokens, roughly 41% of the standard price. The practical split: use Flash for high-volume catalog work and bulk editing, use the standard model for hero assets and high-stakes brand visuals.
It's already inside Microsoft products
This isn't just a developer API release. MAI-Image-2.5 is live on PowerPoint for high-quality image generation and rolling out to OneDrive for precise editing. In PowerPoint, users can generate presentation-ready visuals from prompts. In OneDrive, it powers photo editing like background cleanup and distraction removal while preserving the original scene. That kind of product integration at scale is something OpenAI and Google are still building toward.
Where to use it right now
Today, MAI-Image-2.5 is available for maximum fidelity, and MAI-Image-2.5-Flash for fast, scalable production workloads. There are three access paths:
- MAI Playground , free browser-based sandbox, no API key needed, good for quick evaluation
- Azure AI Foundry , full API access with token-based billing, the production path for teams already on Azure
- OpenRouter , available through the same unified API that OpenRouter's developer community already uses, no new account required if you're already there
Practical use-cases worth knowing
- E-commerce: product shots, background swaps, and consistent catalog edits at scale
- Marketing: on-brand visuals with reliable text rendering for campaigns and decks
- Presentation tooling: PowerPoint-ready diagrams and slide visuals generated from prompts
- Photo editing products: localized edits that preserve the original scene, like the OneDrive integration
- Creative pipelines: generation and editing behind a single API, with identity consistency across a character or product across multiple shots
The fine print
Like all image models, MAI-Image-2.5 can reflect biases in its training data and may produce plausible but inaccurate or misleading visual details. Generated images should be reviewed before use in sensitive contexts, including identity, legal, medical, financial, or news-related workflows. Microsoft has built in layered safety guardrails including prompt and output filtering, but the standard caveats around generative image models apply.
Unlike Microsoft's DALL-E-based tools like Bing Image Creator and Designer, MAI-Image-2.5 is Microsoft's own proprietary model. The MAI team is clearly building a model family it owns end-to-end, and the pace of iteration from MAI-Image-2 to 2.5 suggests this is not a one-off push. With a next-generation GB200 cluster now operational at the MAI lab, the next version is likely already in training.