Artificial Analysis Runs 30 Edits to Expose How Image Models Drift
A 30-edit real estate staging stress test shows how much of the frame each top image model actually touches, and how fast errors compound.
- Artificial Analysis stress-tested four frontier editors with 30 consecutive edits on one photo.
- Ideogram 4.5 and FLUX 3 preserved 95%+ of pixels on small local edits.
- GPT Image 2.5 Sunburst re-renders most of the frame, leaving only ~20% unchanged per edit.
- Nano Banana 2.1 edits locally but the background gradually darkens across turns.
- Sunburst still leads the single-edit leaderboard at 1,197 Elo; Ideogram 4.5 ranks #23.
- Takeaway: pick copy-and-patch models for iterative workflows, full re-render for one-shot generation.
Thirty consecutive edits reveal which image models drift
Artificial Analysis stress-tested four image editors by applying 30 consecutive changes to the same real estate photo. The test exposed how quickly each model altered details outside the requested edit, a failure mode that single-edit leaderboards rarely capture.
Each model received the same sequence of instructions and edited its own previous output. The prompts included lighting a fire, adding a sofa, repainting walls, placing flowers, and changing daylight to twilight. This setup mirrors an interactive workflow in which every new request inherits artifacts from earlier turns.
Thirty turns, one room
- Models tested: Ideogram 4.5, OpenAI GPT Image 2.5 Sunburst, Black Forest Labs FLUX 3, and Google Nano Banana 2.1.
- Starting point: The same real estate photograph for each model.
- Process: Every output became the input for the next edit.
- Primary measurement: The share of the frame that remained effectively unchanged outside the requested modification.
The models fell along a spectrum between local patching and broad re-rendering. Local editors retained source pixels around the requested region. Broader editors regenerated much of the frame, giving them more scope to adjust lighting and style while creating more opportunities for cumulative drift.
| Model | Observed behavior | Long-chain risk |
|---|---|---|
| Ideogram 4.5 | Modified targeted regions while retaining most source pixels. On a small edit such as adding tulips, at least 95% of the frame remained essentially unchanged. | Low pixel drift in the tested sequence. |
| FLUX 3 | Showed similarly local editing behavior, preserving at least 95% of the frame during small changes. | Low pixel drift in the tested sequence. |
| GPT Image 2.5 Sunburst | Re-rendered most of the image during each turn. In the cited comparison, roughly one-fifth of the frame remained unchanged. | Color, texture, and structural changes accumulated across turns. |
| Nano Banana 2.1 | Kept edits relatively local, although surrounding pixels shifted slightly. | The image gradually darkened during the sequence. |
Leaderboard rank misses drift
Ideogram designed version 4.5 to reduce the pixel shifts, color changes, and texture artifacts that accumulate during repeated editing. The test gives that claim a concrete measurement by comparing retained pixels after targeted changes.
Its single-edit rankings tell a different part of the story. According to the reported ranking summary, Ideogram 4.5 debuted at number 23 for image editing and number 34 for text-to-image generation. GPT Image 2.5 Sunburst Max scored 1,197, Grok Imagine Image 2.0 scored 1,155, and Microsoft MAI-Image-2.6 scored 1,150.
Those rankings summarize performance on individual tasks. A 30-turn sequence measures persistence: whether furniture, typography, product details, colors, and composition survive after repeated revisions. Developers building interactive editors need both measurements because a model can produce a strong isolated edit while changing previously approved work.
Match the model to the edit chain
| Model | Likely fit | Relevant capabilities |
|---|---|---|
| Ideogram 4.5 | Product photography, poster revisions, staged interiors, and workflows with protected brand or layout elements. | Up to four reference images, an optional mask, native 2K output, and crop-and-stitch high-resolution editing. |
| GPT Image 2.5 Sunburst | One-pass generation and shorter edit chains that require scene-wide relighting or restyling. | Up to 4K output, as many as 16 reference images, and xhigh and max quality tiers. |
| FLUX 3 | Targeted changes where retaining surrounding pixels is a priority. | Aggressive pixel preservation during the small edits in this test. |
| Nano Banana 2.1 | Local editing sessions with monitoring for brightness and color drift. | Relatively contained changes, with gradual darkening observed over long sequences. |
Local preservation has limits as a quality metric. Copying unchanged pixels can produce a high stability score even when the requested edit fails, while a successful twilight conversion may need to alter lighting across the entire frame. Evaluation should therefore score instruction compliance and protected-region stability separately.
Test the whole session
The benchmark used one photo and one 30-step real estate sequence, so its results do not establish performance across every subject, prompt, resolution, or API configuration. Teams should reproduce the test with their own assets and expected editing patterns before selecting a model.
- Build representative chains: Test the number and type of revisions users commonly request.
- Define protected regions: Track changes to logos, faces, product geometry, text, and approved layout elements.
- Score edit success separately: Verify that each instruction was completed instead of relying only on pixel similarity.
- Monitor cumulative drift: Measure brightness, color, texture, composition, and detail after every turn.
- Budget by session: Include all revisions, retries, output resolutions, and quality tiers in cost and latency estimates.
Ideogram 4.5 costs between $0.008 and $0.22 per image across four quality modes at native 2K resolution. Sunburst is available through the OpenAI API and partners including Picsart and Atlas Cloud. A useful cost comparison multiplies the per-image rate by the expected number of turns and retries rather than treating each generation as an isolated request.
Model selection should follow the expected editing session. Broad re-rendering suits workflows that welcome scene-wide reinterpretation, while local preservation better serves long revision chains with approved elements that must remain fixed. Artificial Analysis found a substantial gap between those behaviors after 30 turns, giving developers a practical benchmark to reproduce against their own workloads.