Google DeepMind's Genie Now Turns Real Street View Into Interactive AI Worlds

Google DeepMind and Google Labs add Street View grounding to Project Genie, letting users build interactive AI worlds anchored in real locations

·
·
  • Google DeepMind and Google Labs added Street View grounding to Project Genie, letting users build interactive AI worlds anchored in real U.S. locations.
  • The feature uses Maps Imagery Grounding and over 280 billion Street View images spanning 20 years to anchor generated environments in reality.
  • Genie generates worlds frame-by-frame (auto-regressively) at 20-24 fps and 720p, maintaining spatial consistency when you turn 360 degrees.
  • Waymo already uses Genie to train self-driving cars on rare, dangerous scenarios — Street View grounding adds real-world geographic specificity to those simulations.
  • Key limitations: no accurate physics, video-game-level fidelity (not photorealistic), U.S. locations only, and sessions limited to a few minutes.
  • Access requires a Google AI Ultra $200/month subscription; the $99.99 tier does not include Project Genie.

Project Genie, Google DeepMind and Google Labs' experimental world-model prototype, just got a significant new capability: Street View grounding. Announced at Google I/O and now rolling out to users, the feature lets you pick a real-world location from Google Maps, apply a creative style, and have Genie generate a navigable, interactive 360-degree environment built on top of that place's actual Street View imagery. This is not a new product , it is a meaningful expansion of an already-live research prototype.

What changed, and what it means

Until this update, Project Genie had been accessible to Google AI Ultra subscribers since January 2026, but was a generator of purely synthetic worlds: you supplied a text prompt, the model invented a coherent environment from scratch, with the starting point always a product of its algorithmic imagination. Street View grounding flips that logic entirely. You now start from a real, geolocated place, photographed and annotated by the Maps team across twenty years of acquisition campaigns, and ask the model to build on top of that substrate a stylized variant, navigable and modifiable in real time.

The integration marks one of the most tangible demonstrations yet of what generative world models can do when paired with a colossal real-world dataset. Project Genie can now draw on more than 280 billion images captured across 110 countries and all seven continents. That archive is a competitive moat no other AI lab can easily replicate.

How it works

A world model , as opposed to a video generator , is a system designed to simulate environments that respond to actions. Rather than producing a fixed clip, it predicts what the world looks like next based on what you do. Genie's environments are described as "auto-regressive" , created frame by frame based on the world description and user actions. This allows for fluid, real-time interaction within the generated world, operating at 20-24 frames per second.

The Street View integration is powered by a technology called Maps Imagery Grounding. This capability is powered by Maps Imagery Grounding, the same technology developers use to create AI visuals with Street View. The workflow is straightforward:

  1. Tap the Maps pin and select a U.S. location
  2. Choose a visual style: "Desert Sands," "Stone Age," "Ocean World," or "B&W film"
  3. Describe a character (an animal, a comic book hero, a claymation monster)
  4. Genie generates an interactive world whose starting point is anchored to that Street View location

What works well, according to Jonathan Herbert, director of Google Maps, is spatial continuity. Turn 360 degrees inside a Genie-generated environment, and the AI remembers what was behind you , it maintains a coherent model of the space rather than regenerating it from scratch with every viewpoint shift. The environments remain largely consistent for several minutes, with memory recalling changes from specific interactions for up to a minute.

The research angle is the real story

Diego Rivas, Group Product Manager at Google DeepMind, described Genie as "our general-purpose world model capable of generating diverse, interactive environments," noting that since launching, it has become a foundational tool for research, enabling agents to learn and reason in complex virtual settings and even helping Waymo simulate hyper-realistic road environments.

The robotics angle is not theoretical. Genie 3 already powers one of Waymo's simulators, where the self-driving car company uses it to train on rare events that would be dangerous or impractical to stage in real life , things like tornadoes or unexpected encounters with elephants on a road. The ability to ground these simulations in actual Street View geography adds another layer of realism.

Parker-Holder illustrated the robotics case with a concrete example: a robot deployed in London, a city that rarely sees direct sun, can use Genie to pre-simulate the rare occasions when sunlight glints off Victorian-era buildings. Without that training data, the robot risks disorientation the first time it encounters the actual glare. Real-world events that are too rare to appear in standard training sets are precisely where generative world models earn their value.

This expansion of Genie's capabilities can provide a virtual environment for AI agents or robots to navigate and interact with the complexities of the real world. That is the deeper ambition here , not just a creative toy, but infrastructure for training physical AI.

Where it falls short

The team is candid about the current limitations. Diego Rivas cautioned that the Street View integration remains experimental. The generated environments look closer to a video game than to a photograph, and the model is not yet physics-aware , in one demonstration, a character ran straight through a row of cacti without consequence.

Parker-Holder acknowledged the gap directly, estimating that interactive world generation trails video generation by roughly six to 12 months in terms of accuracy. The official limitations page lists several known constraints:

  • Physics simulation: Characters can pass through solid objects
  • Real-world accuracy: Genie cannot simulate real-world locations with perfect fidelity
  • Text rendering: Legible text is only reliably generated when included in the input prompt
  • Session duration: The model supports a few minutes of continuous interaction, not extended sessions
  • Geographic coverage: Street View grounding is currently limited to U.S. locations, with additional geographic expansion planned over time.

Who can use it, and how

Project Genie , including the new Street View capability , is gradually rolling out to all eligible Google AI Ultra $200 subscribers globally (18+). Project Genie remains exclusive to that higher tier; subscribers to the newly introduced $99.99 Ultra plan do not have access. You can try it directly at labs.google/fx/projectgenie.

Project Genie is still an experimental research prototype in Google Labs, so Google is working behind the scenes to make the details even sharper and more accurate. The practical use cases span two very different audiences:

  • Creative exploration: Reimagine iconic landmarks , the Golden Gate Bridge underwater, the Fort Worth Stockyards in 1920s black-and-white
  • Agent training: Researchers can use real-world-grounded environments to test and train AI agents in scenarios that mirror actual locations
  • Robotics simulation: Generate rare or dangerous edge-case scenarios grounded in real geography, without physical risk
  • Embodied AI research: Genie is a world model, not just an image generator , it is designed to represent environments that can respond to actions, meaning an agent or user can move through it rather than just observe a static picture.

The competitive picture

This kind of simulation-to-reality pipeline is becoming a critical bottleneck in physical AI. Companies including Nvidia and Cadence have been racing to close the gap between what robots learn inside computers and how they perform once deployed. World Labs, founded by Stanford researcher Fei-Fei Li, released a competing world model called Marble in November 2025, offering commercial access through a freemium structure. Runway launched its own world model in December 2025 with a focus on cinematic applications. None of those competitors arrived with anything approaching 280 billion real-world images collected across two decades. The Street View archive is the specific asset that distinguishes Google's approach: it provides geographic specificity and temporal depth that synthetically generated training data cannot match.

The launch fits within a broader pattern at Google, where the company is steadily threading AI capabilities into products that already have massive user bases. Street View's dataset is a competitive moat that no other AI lab can easily replicate, and connecting it to a generative world model turns a passive mapping tool into something altogether more dynamic. The question is no longer whether world models can generate compelling environments , it is whether they can do so with enough fidelity and physical accuracy to become real training infrastructure. Google is betting that two decades of Street View data gives it a head start on that answer.

Comments

avatar