Reka Releases RIDM to Read Hidden Controls From Any Video
Reka's open-sourced RIDM reads motion from raw video and outputs keyboard-and-mouse style commands, unlocking action-labelled data for world model training.
- Reka released RIDM, a compact inverse dynamics model that labels raw video with WASD keys, yaw and pitch.
- Trained entirely on gameplay footage at 1280x720 and 48 fps with engine-provided ground truth.
- Flow variant uses frozen RAFT-small, 2.8M params total; pixel variant uses IMPALA trunk at 9.8M params.
- Flow model hits 85.9% on real video action classification versus 30.8% for the pixel model.
- Weights open-sourced under Apache 2.0 on Hugging Face, safetensors format.
- Fails on static frames, zoom effects, and stabilised action-camera footage where flow signal is weak.
Reka releases RIDM to infer controls from video
Interactive world models need video aligned with the low-level controls that caused each visual change, yet ordinary footage rarely includes those logs. Reka has released the Reka Inverse Dynamics Model (RIDM) and published the weights under Apache 2.0. Developers can use the compact model to generate pseudo-labels for existing gameplay, YouTube videos, and action-camera footage.
Given a short first-person clip, RIDM estimates probabilities for discrete controls, including W, A, S, D, and Shift as a walking modifier. It also predicts continuous yaw and pitch for camera turns and tilts. Reka built the RIDM training set entirely from gameplay, then evaluated its transfer to real-world footage. The output schema targets first-person navigation and camera control.
Action logs are the bottleneck
A world model learns to predict how an environment changes after an action, allowing a training system to roll out possible futures for tasks such as folding cloth or squeezing a towel. Those interactions are difficult to reproduce faithfully in hand-built simulators. Training an interactive model requires each visual transition to be paired with the action that caused it, while most recorded video supplies only images and manual labeling cannot recover exact hidden inputs.
An inverse dynamics model infers the likely action from an observed visual transition. Running it across unlabeled footage produces a sequence of model-generated pseudo-actions that can be aligned with the source frames and passed to a world-model trainer.
Optical flow crosses the domain gap
Reka tested two compact architectures. The pixel model feeds raw frames through the IMPALA convolutional trunk used in OpenAI’s VPT work. The flow model uses a frozen RAFT-small network to convert frame pairs into optical flow before a trainable transformer processes the motion.
| Model | Representation | Parameters | Released weights |
|---|---|---|---|
| Pixel | Raw frames through IMPALA | 9.8 million, all trainable | About 39 MB |
| Flow | RAFT-small optical flow | 2.8 million total, 1.8 million trainable | About 7.2 MB |
Optical flow records where image points move between frames while discarding much of their color and texture. A game corridor and a real street can therefore produce similar flow fields under the same camera movement, despite looking different at the pixel level. Removing appearance cues before prediction reduces overfitting to game graphics.
Real footage separates the models
Reka’s published classification results measure whether each model identifies forward motion, a left turn, or a right turn. A separate two-class test measures turn direction.
| Dataset | Task | Flow model | Pixel model |
|---|---|---|---|
| Real video, 78 clips | Forward, left turn, or right turn | 85.9% | 30.8% |
| Real video, 47 turns | Left or right turn | 91.5% | 51.1% |
| GoPro, 72 clips | Forward, left turn, or right turn | 75.0% | 58.3% |
| Gaming footage | Forward, left turn, or right turn | 84.5% | 77.5% |
Across Reka’s full six-configuration comparison, the flow model leads in five. Its largest listed margin is 55.1 percentage points on the 78-clip real-video set. The pixel model’s sole lead comes from GoPro turn direction, where it classifies one additional clip correctly.
Motion-balanced clips shape training
Reka recorded gameplay at 1,280 by 720 pixels and 48 frames per second, logging 11 key commands, mouse movement, and exact camera angles from the game engine for every frame. Forward motion dominated the raw recordings, so the team sampled five motion-specific pools:
- Forward: about 105,000 clips with
Wheld and minimal camera motion. - Turn: about 105,000 clips with no directional keys and at least 10 degrees of yaw.
- Lateral or reverse: about 53,000 clips using
A,D, orS. - Tilt: about 1,500 clips, making it the rarest sampled motion.
- Lateral plus turn: 50,000 clips added in the final stage to separate strafing from camera rotation.
Failures follow weak motion signals
RIDM’s reported errors cluster around footage where optical flow is absent, removed, or ambiguous:
- Low signal: Static or very dark frames can trigger false turn predictions.
- Conflicting motion: A sideways step combined with an opposite yaw turn can confuse translation and rotation.
- Scale changes: Zooming in or out can resemble forward or backward movement.
- Stabilization: Action-camera stabilization can remove rotation before RIDM processes the clip.
Because these cases remove or distort the motion field, production pipelines should retain key probabilities and review low-confidence or known-risk segments before using the predictions for training.
Open weights carry a RAFT caveat
The Hugging Face release stores weights in safetensors format, with a config.json file and an inference.py script in each model folder. The flow artifact is about 7.2 MB, while the pixel artifact is about 39 MB.
Reka licenses its RIDM artifacts under Apache 2.0. The flow variant also depends on Torchvision’s pretrained RAFT-small weights, whose training-dataset terms may restrict commercial use. Commercial deployments require a review of those terms, while the pixel model avoids the RAFT dependency.
Turning clips into pseudo-labels
Developers can integrate RIDM as an annotation stage by preserving clip alignment and treating each prediction as an uncertain pseudo-label:
- Segment first-person footage into the clip windows defined by the selected model’s configuration.
- Run inference and retain the discrete key probabilities plus continuous yaw and pitch values.
- Attach each output to its corresponding clip window.
- Filter or review dark, static, stabilized, zoom-heavy, and conflicting-motion clips.
- Feed the accepted video-action pairs into the downstream training pipeline.
World-model and controllable-video trainers can consume the aligned labels as conditioning data, while navigation-policy pipelines can use them for pretraining or data triage. Embodied systems still need a separate mapping from RIDM’s first-person controls to the target robot’s actuator space.