Google's Nano Banana 2.1 Beats Pro on Image Editing at Flash Prices
Google's updated workhorse image model keeps Flash-level speed while beating Nano Banana 2 and even Pro on most editing benchmarks.
- Google released Nano Banana 2.1, an update to its Flash-tier image generation and editing model.
- Built on Gemini 3.6 Flash, available as
gemini-nano-banana-2.1with configurable Thinking (minimal, medium, high). - Beats Nano Banana 2 and even Nano Banana Pro on most human-eval Elo editing benchmarks.
- Supports up to 14 reference images, 4K output, and Google Search grounding for factual infographics.
- Rolling out in Gemini app, AI Studio, API, Flow, Stitch, Google Ads, and Search AI Mode.
- Known weak spots: small text, spatial left/right reasoning, and occasional pose leakage during edits.
Google says Nano Banana 2.1 beats Pro on key image tests
Google has released Nano Banana 2.1, a point update to its speed- and cost-optimized image model. API access is available under gemini-nano-banana-2.1, while a broader rollout is underway across Google’s developer and consumer products.
Google positions the release as a quality upgrade that retains Nano Banana 2’s Flash-tier pricing and latency. Nano Banana 2 corresponds to Gemini 3.1 Flash Image; according to the release documentation, version 2.1 belongs to the Gemini 3 series and builds on Gemini 3.6 Flash.
Sharper edits at Flash speed
The update targets four areas that directly affect image-generation and editing workflows:
- Visual composition: More polished layouts, assets, and scene construction.
- Mask-based editing: More precise changes to selected regions without regenerating the full image.
- Subject consistency: Better preservation of characters and objects across edits and related images.
- Rendering quality: More natural results with fewer visible artifacts, including improved text rendering.
The model also expands the controls and inputs available to developers:
| Capability | Nano Banana 2.1 support |
|---|---|
| Output resolution | 1K, 2K, and 4K, with improved realism and fewer tiling artifacts in wide 2K and 4K images |
| Reference images | Up to 14, including consistency for up to four characters and fidelity for up to 10 objects |
| Grounding | Google Web and Image Search results can inform generated content |
| Thinking levels | Minimal, medium, and high; medium is the default |
| Context window | 131,072 tokens |
| Accepted inputs | Text, images, video, and PDFs |
Google’s tests favor 2.1
Google’s internal side-by-side evaluation asks human raters to choose between model outputs, then converts those preferences into Elo scores. Higher Elo indicates stronger relative preference within the test; point differences do not translate directly into percentage-point quality gains.
| Reported metric | Nano Banana 2 | Nano Banana 2.1 with Thinking | Nano Banana Pro |
|---|---|---|---|
| Overall preference, Elo | 981 | 1050 | 935 |
| Multi-character consistency, Elo | Not reported | 1106 | 1011 |
| Infographic factuality | 0.179 | 0.521 | Not reported |
The infographic factuality score is about 2.9 times the score reported for Nano Banana 2. Search grounding may contribute to that improvement, although the published results do not isolate the effect of grounding from model, training, or reasoning changes.
These results make 2.1 competitive with Pro on the tested editing tasks, especially when Thinking is enabled. The benchmarks remain vendor-reported, use selected prompts and evaluation criteria, and may not predict performance on a team’s own assets.
A familiar API surface
The Python call through the Interactions API requires a model ID and prompt:
from google import genai
import base64
client = genai.Client()
interaction = client.interactions.create(
model="gemini-nano-banana-2.1",
input="A cyberpunk ramen stand at night, neon reflections on wet pavement",
)
with open("out.png", "wb") as f:
f.write(base64.b64decode(interaction.output_image.data))Because the example omits a Thinking setting, it uses the medium default. Minimal suits latency-sensitive generation, while high targets infographics, multi-subject scenes, and other prompts that require more planning. Google says pricing and expected latency remain aligned with Nano Banana 2; current prices, quotas, and configuration fields are available in the model documentation.
The rollout covers the Gemini app, AI Mode in Google Search, Google AI Studio, Google Flow, Stitch, Google Ads, and the Gemini Enterprise Platform. Product availability may vary by account, region, and rollout stage.
Failure modes to test
Google identifies several limitations that can affect production use:
- Small text can appear blurry at 1K, while long paragraphs remain difficult to render accurately.
- Character appearance can drift between reference and generated images.
- Masked and doodle-based edits may follow only part of an instruction.
- Source poses can persist when the requested edit requires a structural change.
- Spatial directions such as left and right can be confused.
- Three-dimensional reasoning and factual accuracy remain imperfect.
Google recommends 2K or 4K output when small text matters. Production workflows should still validate spelling, labels, factual claims, identity consistency, and spatial relationships before publishing generated assets.
How to route real workloads
Nano Banana 2.1 gives teams a faster default candidate for editing, grounded infographics, product imagery, and character-driven campaigns. The appropriate Thinking level and output resolution depend on the workload:
| Workload | Suggested starting point | Primary check |
|---|---|---|
| High-volume drafts | 2.1 with minimal Thinking | Confirm that lower latency preserves required quality |
| General image editing | 2.1 with medium Thinking | Check prompt adherence and mask boundaries |
| Infographics or complex scenes | 2.1 with high Thinking | Verify facts, labels, and subject placement |
| Multi-reference compositions | 2.1 with medium or high Thinking | Measure character and object consistency |
| Assets dominated by small text | 2.1 at 2K or 4K, compared with Pro | Inspect text at final display size |
Teams evaluating a migration should compare both models on representative prompts, final-size assets, latency, and total generation cost. Google’s published scores support using gemini-nano-banana-2.1 as the first candidate for many editing workflows, while workload-specific testing determines whether Pro still earns a place in the routing stack.