Google Builds Nano Banana 2 Into Earth to Reimagine Any Location
Google Earth now lets you generate AI images of any location on Earth using Nano Banana 2, grounded in real satellite and 3D data.

- Google has integrated Nano Banana 2 image generation into Google Earth on the web, live globally and free today.
- Users zoom into any location, tap "create image," and type a prompt; the model generates images grounded in real satellite and 3D data.
- Nano Banana 2 (Gemini 3.1 Flash Image) combines Pro-level quality with Flash speed, supporting up to 4K resolution and 5-character consistency.
- Use cases include historical visualization, real estate rendering, urban planning, location infographics, and creative reimagining.
- The feature is currently web-only, excluding the Google Earth mobile app despite 500M+ Play Store downloads.
- All generated images are watermarked with SynthID and paired with C2PA Content Credentials for AI provenance tracking.
Google has integrated Nano Banana 2 image generation into Google Earth on the web, letting anyone zoom into any spot on the globe, type a prompt, and get a photorealistic AI-generated image grounded in actual satellite, aerial, and 3D map data. It is live globally, for free, right now.
How it works
The workflow is straightforward: open Google Earth on the web, zoom into a location, tap "create image," and type what you want to see. The model uses the real geometry and visual context of that location as a base, then layers your prompt on top. A prompt like "add a modern lakefront cabin" produces a rendering of that specific lot, at that specific angle, with the surrounding landscape intact rather than a generic cabin dropped into a void.
The model behind it
Google DeepMind built Nano Banana 2 on the Gemini 3.1 Flash Image architecture, combining the quality of Nano Banana Pro with the speed of Gemini Flash. The previous Nano Banana models were based on the 3.0 branch.
The architectural difference from traditional diffusion models matters for this use case. Rather than treating a prompt as a bag of weighted keywords, Nano Banana 2 reasons about composition, lighting, and spatial relationships before rendering anything. It reads intent and context the way a language model would. Gemini 3.1's broader world knowledge also feeds the system, giving it enough factual grounding to render objects accurately and generate location-specific infographics rather than plausible-looking fabrications.