Runway's Solaris Ditches Code to Build Apps as Live Video
Runway's Solaris generates application interfaces pixel by pixel in real time, skipping code entirely and letting a world model render each frame as you click.
- Runway unveiled Solaris, an Interface World Model that renders app UIs frame by frame with no code.
- Built on Gen-4.5 video model, pairs an LLM for reasoning with a world model for rendering at 720p.
- In a 7,500-judgment user study, beat Claude Opus 5 coded interfaces 71% to 21% on natural behavior.
- Uses autoregressive generation plus distillation to hit interactive, sub-second per-frame latency.
- Weak at legible text, long-session coherence, factual trust, and accessibility like screen readers.
- Available by early-access request only; aimed at retail, tutorials, product visualization, and agent training data.
Runway Solaris generates interfaces as live video
Runway has unveiled Solaris, a research preview that generates an interactive interface as a stream of video frames. During a Runway AI Summit keynote, Kamil Sindi described it as the first system in a new family called Interface World Models. User actions such as clicks, drags, and typed input condition each new frame, with no interface source code, DOM, or component tree behind the generated view.
Runway’s longer-term proposal extends this approach to applications and operating systems generated continuously by neural networks. The current preview has a narrower scope: it demonstrates model-rendered interactions while leaving application logic, data storage, authentication, payments, and other production infrastructure unresolved.
Pixels become the runtime
Conventional interfaces begin with a design, encode that design in components and application logic, and let a browser or operating system render the result. Solaris receives a starting image and a history of user input, then predicts the visual response directly. A click becomes a conditioning signal for the next frame; no event handler exists within the generated interface.
The company argues that implementation constrains a design to behaviors engineers define in advance. Solaris can generate transitions, motion, and visual states that were never represented as reusable components. Code also supplies guarantees that pixels alone cannot provide, including semantic markup, accessible controls, deterministic state transitions, form behavior, and integration with back-end systems.
The model behind the pixels
Solaris builds on Gen-4.5, Runway’s video diffusion model, with adaptations for user input and interactive latency. Its architecture separates high-level decisions from frame generation:
- A language model plans the response. It interprets requests such as changing a couch’s color, determines whether an action modifies the current scene or opens a new one, and produces instructions for the world model.
- A world model renders the response. It consumes clicks, drags, typed input, prior frames, and generated instructions, then predicts the next visual state.
Runway describes three techniques used to reduce diffusion latency and limit quality loss during longer sessions:
- Autoregressive generation lets each frame depend on previous frames without generating an entire clip at once.
- Diffusion distillation compresses a denoising process that normally requires dozens of steps into a smaller number of steps.
- Training on generated outputs exposes the faster student model to its own prior frames, helping it recover from errors that accumulate over time.
The published targets include 720p output and latency below 500 milliseconds per generated frame. Runway has not provided independent latency tests, hardware requirements, throughput figures, or public benchmarks for extended sessions.
Runway’s benchmark claims
Runway reports two evaluations of the approach. The first asked GPT-4o, Gemini 2.5 Pro, and Claude Fable 5 to reconstruct 30 interface screenshots as code. The resulting pages were compared with the source images using SSIM, which measures structural similarity, and DINOv3 feature similarity, which estimates correspondence between visual features. Scores declined as the interfaces became more visually complex, with the largest losses on natural imagery.
The second evaluation compared Solaris with websites generated by Claude Opus 5 from the same starting image and interaction request. Runway reports about 7,500 pairwise judgments from 250 participants:
| Evaluation criterion | Preferred Solaris | Preferred Claude |
|---|---|---|
| Following instructions | 61% | 24% |
| Behaving naturally | 71% | 21% |
The reported shares do not total 100%, and the summary does not explain the remaining responses. These results come from Runway’s own study and have not been independently replicated. Prompt design, implementation quality, hardware, response time, and task selection could materially affect the comparison.
Text, drift, and missing semantics
Runway positions Solaris as a research preview suited to ambient motion, direct visual manipulation, and scene transitions. Demonstrations include placing clothing on a photo, rearranging furniture, and moving through a visual tutorial that adapts to user actions. Several limitations currently block general production use:
- Typography: Video models still struggle to maintain legible, stable text across frames. Runway proposes using image models to render text-heavy views during pauses.
- Grounding: A visually plausible response may conflict with product data, instructions, inventory, or other authoritative sources. Reliable grounding remains an open research area.
- Session drift: Generated state can lose semantic coherence as interactions accumulate.
- Determinism: Identical actions can yield different frames, complicating regression tests, debugging, replay, and audit logs.
- Accessibility: The preview has no disclosed screen-reader support or semantic accessibility tree, and unstable text further limits access.
- Application plumbing: Public materials do not define how generated controls connect to authentication, databases, analytics, payments, validation, or secure API calls.
Production systems would need an explicit boundary between generated presentation and authoritative state. Prices, permissions, inventory, financial transactions, and safety rules require validation outside the visual model, regardless of what Solaris renders.
Three UI paths, three artifacts
Solaris occupies a different technical category from code-generation products such as Vercel’s v0, Lovable, and Bolt. Those tools generate source files that standard browsers execute. Solaris generates the visible interaction itself.
| Capability | Conventional application | AI code generator | Solaris |
|---|---|---|---|
| Primary artifact | Components, styles, and application logic | Generated source code | Generated video frames |
| Runtime | Browser or native platform | Browser or native platform | World model |
| Interaction handling | Defined event and state logic | Generated event and state logic | Predicted visual response |
| Accessibility | Platform semantics and accessibility APIs | Depends on generated code | No support disclosed |
| Testing | Standard unit, integration, and browser tools | Standard tools after generation | Deterministic testing remains unresolved |
| Back-end integration | Application APIs and services | Generated or manually added integrations | No public integration model |
The project extends Runway’s work on world models, systems that maintain a representation of an environment and predict its next state. The company released GWM-1 for robotics and spatial exploration in December 2025. Solaris applies the same premise to interface environments, treating a shopping cart or product configurator as a world whose state can be visually predicted.
Synthetic interfaces for agent training
Runway also proposes Solaris as a training environment for computer-using agents. Current agents often overfit to familiar layouts and fail when buttons, menus, or workflows move. A generative interface model could produce a large supply of novel layouts and interaction sequences, giving agents broader practice without requiring developers to build each environment manually.
That use case would still require controllable tasks, known success criteria, reproducible state, and reliable labels. Solaris could supply visual diversity, while a separate system would need to define goals and verify whether an agent completed them correctly.
Early access, no API
Access currently requires a request through the research page. Runway has announced no public API, pricing, open weights, deployment model, or production service-level commitments. Teams building retail configurators, product visualizers, adaptive tutorials, and agent-training environments align most closely with the demonstrations shown so far.
Production adoption will depend on stable text, lower and measurable latency, longer coherent sessions, accessibility support, deterministic controls, and secure connections to application data. Solaris currently offers a research direction for model-generated interfaces rather than a deployable replacement for the web application stack.