Odyssey Releases Odyssey-3, a World Model That Teaches Robots Physics Through Video
Odyssey-3 is a real-time foundation world model that generates interactive 3D environments from prompts and tops the Physics-IQ benchmark at 66.1.
- Odyssey released Odyssey-3, a real-time interactive foundation world model generated from text prompts.
- Odyssey-3 Pro scores 66.1 on Physics-IQ Verified video-to-video, the highest reported score.
- Ranks first in 3 of 4 WorldMark categories, trailing only in first-person real environments.
- Frozen backbone drives a car in India after just 20 hours of driving data.
- Autoregressive diffusion transformer distilled to few-step sampling for real-time interactive generation.
- Free research preview live at experience.odyssey.systems, API access by request.
Odyssey-3 links interactive video with physical control
Odyssey has released Odyssey-3, a foundation world model that generates navigable environments from text and predicts how each scene changes as a person or software agent acts. The free research preview is available now, while developers working on robots and vehicles can request API access.
At its core, Odyssey-3 is an autoregressive diffusion transformer: it generates visual states through iterative denoising, then conditions each new state on previous observations and control inputs. Odyssey designed the same pretrained backbone for three workloads: interactive simulation, physical-machine control, and agent training. Reusing that backbone could reduce the task-specific data and compute required for each robot, vehicle, or simulated environment.
A physics lead with a compute caveat
Physics-IQ Verified evaluates whether a model can continue recordings of real experiments while preserving behavior across fluid dynamics, optics, solid mechanics, magnetism, and thermodynamics. Odyssey-3 Pro achieved the highest reported video-to-video score, according to the launch results.
| Benchmark | Result | Context |
|---|---|---|
| Physics-IQ Verified, video-to-video | 66.1 | Odyssey-3 Pro at 720p using best-of-eight sampling; reported ahead of Black Forest Labs’ FLUX 3 and NVIDIA Cosmos3 variants. |
| Physics-IQ Verified, image-to-video | 54.7 | Odyssey-3 Pro using best-of-eight sampling. |
| WorldMark, first-person stylized | First | Ranking based on the mean of 13 reported metrics using WorldMark’s captions. |
| WorldMark, third-person real and stylized | First | Odyssey-3 led both third-person splits. |
| WorldMark, first-person real | 80.6, third | Behind Lyra 2.0 at 84.4 and AlayaWorld at 83.0. |
The 66.1 Physics-IQ result uses the strongest outcome from eight generated samples, which requires more compute than a single generation. Odyssey also reports that the base model improves the tradeoff between physical accuracy and generation cost, although the launch materials provide charted generation costs rather than API pricing or hardware-specific latency.
Three data streams teach cause and effect
Odyssey trained the model on three complementary data sources, each covering a different part of interactive world modeling:
- Internet video supplies broad visual coverage, paired with time-localized event annotations checked against a defined schema.
- Gameplay recordings pair frames with synchronized keyboard and mouse inputs, giving the model explicit action-to-outcome examples.
- Rigid-body simulations provide controlled interactions with captions and metadata, making cause and effect easier to isolate.
Streaming diffusion under the hood
The base architecture is a multi-step video diffusion transformer with temporal rotary positional embeddings, or RoPE, which encode the order of frames. Causal masking prevents the model from attending to future frames, allowing it to predict the next state from past observations. During training, teacher forcing supplies the correct preceding observations; during generation, the model continues from its own output while accepting new prompts and controls.
Odyssey compresses the original diffusion sampler into a few-step model to reduce interaction latency. Post-training combines distribution-matching distillation, which teaches a faster model to reproduce the slower model’s output distribution, with generative adversarial discriminators and reinforcement learning. The company has not published an exact frame rate, context length, or hardware profile for the research preview.
One backbone, three machines
Odyssey adapts the frozen visual backbone to physical systems by training an action decoder or policy on paired observations and actions. Keeping the backbone frozen preserves its pretrained weights while the smaller policy learns how to control a particular machine.
- Robot arms. Odyssey reports that policies trained on tens of hours of demonstrations completed manipulation tasks and recovered from failures absent from the demonstrations. Examples include reorienting a gripper after a missed grasp and retrieving an object dropped in an unusual position.
- Humanoids. Flexion built humanoid policies on Odyssey-3. The resulting systems outperformed the tested vision-language-action baselines under environmental changes and continued operating under lighting changes that disrupted those baselines.
- Vehicles. Odyssey trained a driving policy on 20 hours of road data from India while leaving the backbone unchanged. The policy uses its visual representations to predict waypoints and drive in closed loop, meaning each prediction affects the next observation and action.
Freezing the backbone can lower adaptation costs because each embodiment needs a smaller policy rather than a newly trained visual model. The public evidence currently consists of Odyssey’s demonstrations and reported evaluations, so independent testing still needs to establish reliability, safety, and performance across unfamiliar machines and environments.
Agents learn inside generated scenes
For agent training, Odyssey-3 serves as the environment that produces observations and responds to actions. In the task-completion demonstration, an agent receives a natural-language goal, observes the generated scene, and takes actions until it reaches the requested outcome.
Generated environments can produce varied rollouts without requiring a physical robot or a hand-authored asset pipeline for every scene. They also carry the world model’s errors into training: inaccurate dynamics, weak long-term memory, or missing edge cases can teach a policy the wrong behavior. Odyssey has not reported task-success rates, agent-training costs, or sim-to-real transfer results for this use case.
Evidence gaps developers should track
- Long-horizon consistency: WorldMark’s world-memory measure is folded into a 13-metric mean, leaving no detailed breakdown of how long scenes remain coherent.
- First-person realism: Odyssey-3 ranks third on WorldMark’s first-person real split, behind Lyra 2.0 and AlayaWorld.
- Runtime requirements: Exact latency, throughput, memory use, context limits, supported control schemas, and deployment hardware remain unspecified.
- Sampling cost: The leading Physics-IQ result requires eight samples at 720p, raising compute use relative to a single interactive generation.
- Physical-system validation: The robot, humanoid, and vehicle results need independent replication, longer deployments, and safety evaluation.
The bet is reusable world knowledge
Odyssey-3’s developer proposition centers on reuse: pretrain visual dynamics across broad video and simulation data, freeze that shared backbone, then train a smaller decoder for each robot, vehicle, or agent. If independent evaluations reproduce the reported transfer results, the approach could reduce specialized data collection and shorten the path from a general visual model to a working control policy.
Preview now, API by request
Developers can try the research preview and request API access for physical systems. Odyssey’s related work includes the Agora-2 multi-agent world model and PROWL-2, a reinforcement-learning framework that trains multiple agents and their world model within one loop.