Reka's Rho-1 Merges Reasoning, Video, and Robot Control Into One Model
Reka's 19B parameter Rho-1 unifies text, images, video, and robot actions in one network, streaming steerable video in real time.
- Rho-1 is a 19B omni model handling text, images, video and robot actions in one network.
- Trained from scratch on 320 H100s in roughly three months, a fraction of frontier compute.
- Distilled Rho-1 Flash renders 5.3-second video clips in about one second via 8-step denoising.
- Symmetric two-stream transformer with shared attention; base model uses 99 denoising passes.
- Streams continuous video that can be steered mid-rollout without cuts, enabling live world simulation.
- Limitations: 672x384 resolution cap, long-horizon structural drift, brittle video grounding and edits.
Rho-1 puts reasoning, video, and robot control in one model
Reka has released a research preview of Rho-1, a 19-billion-parameter model that processes and generates text, images, video, and robot actions within one checkpoint. Its shared context lets generated media, reasoning tokens, and control signals inform subsequent steps in the same session.
Multimodal products often connect several specialized services, such as a language model for planning, a diffusion model for video, a detector for object localization, and a policy model for robot control. Those boundaries require developers to serialize state, move data between services, and preserve context across incompatible representations. Rho-1 moves those operations into one attention context.
One model holds the whole session
The clearest demo begins with a request to draw a lighthouse scene. The user then asks Rho-1 to locate the lighthouse with a bounding box, animate the image as a drone fly-in, add a snowstorm while preserving the camera movement, and explain the differences between the resulting videos. One model maintains the session state throughout the sequence.
Reka describes the architecture as symmetric because it can consume and produce each supported representation. Conventional vision-language models typically accept images and video but return text, while video generators turn multimodal instructions into pixels. Rho-1 keeps generated content available to later operations inside the model’s context, removing the need for an external model handoff.
Two streams share the context
Each transformer block divides computation between two expert weight streams that use a shared attention operation:
- Understanding stream: Processes language, internal reasoning tokens, and visual information. A next-token head emits discrete outputs.
- Generation stream: Uses a flow-matching head to denoise continuous latent representations into images and video.
- Shared attention state: Instructions, previous visual content, reasoning tokens, and motor commands remain available through a common KV cache.
- Internal routing: A special token signals when a response requires visual generation, allowing the generation stream to continue from the accumulated state.
The model carries information in two formats. Discrete tokens represent text, symbolic reasoning, and high-level commands. Continuous representations carry image latents, video frames, robot actions, and proprioception, meaning measurements such as joint position and movement. Preserving physical signals as continuous values avoids the precision loss introduced when they are forced into a discrete vocabulary.
Flash cuts 99 denoising steps to eight
Reka reports the following generation speeds for the base model and a distilled version called Rho-1 Flash:
| Variant | Denoising steps | Reported performance |
|---|---|---|
| Rho-1 | 99 | Median generation rate of 0.79× real time, with streaming output beginning after roughly six seconds |
| Rho-1 Flash | 8 | A 5.3-second video generated in about one second |
| Rho-1 Flash edits | 8 | Roughly one-second clips re-rendered in about 1.1 seconds |
The streaming interface can extend a clip from its previous latent state and incorporate instructions during generation. In one aerial demo, commands to "bank left" and "bank right" arrive half a second into the rollout. The resulting branches retain the same opening frames before following different trajectories.
These figures come from Reka’s preview and cannot yet be reproduced with public weights or an API. Hardware configuration, batching, output settings, and compilation choices can materially change video-generation latency.
Robot actions join the visual sequence
The shared sequence also lets Rho-1 predict visual observations and emit robot actions from the same checkpoint. A policy can generate a likely future camera view alongside motor commands, giving the model a short visual rollout associated with its planned movement.
On a LIBERO simulation task, Rho-1 emits seven action channels together with a predicted wrist-camera view. The architecture removes the service boundary between an external world model and a separate control policy, although the preview does not establish how well the approach transfers from simulation to physical robots.
Robot training data remains scarce because teleoperated trajectories are expensive to collect and label. Reka pairs Rho-1 with its inverse-dynamics model, which estimates the control signals that could have produced an observed video. This process can convert suitable unlabeled footage into inferred action data, with usefulness limited by domain fit and the accuracy of those inferred controls.
320 H100s train the shared backbone
The checkpoint used for the published results was trained from scratch on 320 Nvidia H100 GPUs for three months. That training run covered language, visual understanding, media generation, and control within one backbone.
Two concurrent objectives update the model:
- Next-token prediction trains discrete sequences such as text and symbolic commands.
- Flow matching trains continuous generation for images, video, and action trajectories.
A shared backbone allows training signals from multiple tasks to modify common representations. Reka’s central research claim is that improvements in reasoning, visual generation, and physical control can reinforce one another. The preview does not include public checkpoints or sufficient ablation results to measure that effect independently.
Drift sets the current ceiling
Reka identifies several limitations that block routine production use:
- Long-horizon drift: A 30-second stream can preserve texture and fine detail while changing the room into a structurally inconsistent layout.
- Temporal grounding: Bounding boxes work on static images but do not track objects reliably through video.
- Editing stability: Targeted edits succeed on selected prompts but remain brittle across broader inputs.
- Resolution: Native video rollouts are capped at 672 × 384 pixels.
The published demonstrations show interactive visual continuation, but they do not establish physically calibrated simulation, dependable long-horizon planning, or safe closed-loop robot control. A unified model also concentrates failures: an error in its shared state can affect reasoning, generation, and action output together.
Fewer handoffs, gated access
For application developers, the architecture could simplify state management across editing, analysis, simulation, and control. Every operation can reference the same session context, reducing conversion layers and network calls between specialist models. Deployment would still require substantial accelerator memory and serving capacity for a 19-billion-parameter model that generates video.
Potential uses include branchable robotics simulations, interactive video tools, and policies that predict visual outcomes before issuing actions. Their practical value depends on longer coherent rollouts, stronger temporal grounding, higher resolution, and evidence that the generated futures correspond to real system dynamics.
Rho-1 remains a research preview available through collaboration on Reka Cloud. Reka has not released public weights or a public API, so developers cannot yet integrate the model directly or independently verify its performance claims.