Odyssey's PROWL-2 Fixes Its Own Simulator Errors While Training AI Teams

Odyssey's new framework pairs agents with a self-repairing world model, lifting StarCraft win rates by up to 91% and cracking the hardest robot coordination tasks.

·
·
Read5 min
TypeNews
SubtopicVla Models
  • Odyssey released PROWL-2, a framework where agents and their world model co-train via separate curricula.
  • Reports up to 91% relative gain over baseline across all nine SMACv2 StarCraft scenarios.
  • Lifts Gate-3 robot coordination from 29.8% to 70.4%, Shepherd-Hard from 7.3% to 28.2%.
  • A novel fidelity gate blocks hallucinated rollouts from policy training and routes them to world-model repair.
  • Framework is agnostic to the policy algorithm and world model, so it drops onto existing systems.
  • Full paper details the dual curriculum and SMACv2 plus MQE evaluations.

PROWL-2 repairs world models while multi-agent policies train

Odyssey has introduced PROWL-2, a reinforcement-learning framework that trains a team of agents while identifying and repairing errors in the world model that generates their simulated experience. The policy curriculum prioritizes scenarios where agents can still improve, while a separate exploration policy searches for states the world model predicts poorly.

The accompanying research paper reports improvements over its baseline world-model learner across nine StarCraft Multi-Agent Challenge v2 scenarios, with relative gains reaching 91%. On multi-robot control benchmarks, Gate-3 success rises from 29.8% to 70.4%, while Shepherd-Hard improves from 7.3% to 28.2%.

Shared imagination compounds errors

Model-based reinforcement learning uses a learned simulator, or world model, to predict how an environment will respond to an agent’s actions. Policies can practice through these simulated sequences, known as imagined rollouts, while using fewer interactions with the real environment.

Prediction errors can corrupt that training. A policy may learn to exploit behavior that works inside the model and fails when transferred back to the real environment. Longer rollouts amplify small errors as each inaccurate prediction becomes input for the next step.

Multi-agent systems add another source of drift because every policy changes the environment experienced by the others. As the team learns, the world model must follow shifting interactions among several agents. A mistaken prediction about one agent can alter the simulated trajectory of the entire group.

One loop, three controls

PROWL-2 extends Unsupervised Environment Design, which selects or generates training scenarios according to their expected learning value. A start state is a saved environment configuration from which an imagined rollout begins. The framework coordinates those states through three components:

Diagram of the PROWL-2 policy curriculum, world-model curriculum, and fidelity gate
PROWL-2 connects separate policy and world-model curricula through a fidelity gate.
  1. Task-policy curriculum: Policy agents train in imagined rollouts seeded from start states ranked by learning potential. The curriculum promotes harder states as performance improves.
  2. World-model curriculum: A separate “developer” agent explores the real environment for trajectories the current model predicts poorly. Those failures become fine-tuning data, extending the approach introduced in PROWL-1.
  3. Fidelity gate: The system replays recorded actions from high-priority start states through the world model, then compares each imagined trajectory with its stored real counterpart. States that exceed the permitted error threshold leave the policy curriculum and enter the model-repair queue.

States return to policy training after the repaired model predicts their withheld trajectories accurately enough. Known model failures therefore supply repair examples and remain unavailable for policy updates until validation succeeds.

Harder coordination widens the gap

StarCraft Multi-Agent Challenge v2 procedurally generates battles, reducing the value of memorizing fixed scenarios. The evaluation covers symmetric 5v5 and 10v10 fights, plus asymmetric 10v11 battles where the trained team is outnumbered. Rollout visualizations in the paper show PROWL-2 agents concentrating fire, closing ranks, and using diversions, while baseline teams more often fragment and split their attacks.

Selected SMACv2 win rates over 40 held-out rollouts
Scenario PROWL-2 Baseline
Terran 5v5 87.5% 7.5%
Terran 10v11 55% 2.5%
Protoss 10v11 35% 0%
Zerg 10v10 90% 0%

Each win in the 40-rollout sample represents 2.5 percentage points, so the reported rates remain point estimates from a modest evaluation set. The table shows four of the paper’s nine StarCraft scenarios.

The Multi-Agent Quadruped Environment tests continuous control and physical coordination. At 600,000 training steps on Gate-3, the baseline solves none of 20 evaluation scenarios, while PROWL-2 solves more than 20%. By 1.5 million steps, their learning curves reach roughly 25% and 70%, respectively.

Reported success rates on difficult quadruped tasks
Task PROWL-2 Baseline
Gate-3 70.4% 29.8%
Shepherd-Hard 28.2% 7.3%

Ground truth remains in the loop

The fidelity gate depends on real-environment trajectories to detect and repair simulator errors. PROWL-2 can reduce the number of interactions spent on policy training while retaining ongoing costs for data collection, trajectory storage, validation, and model fine-tuning.

The reported experiments use game and continuous-control benchmarks with defined states, actions, and success criteria. Physical deployments would add sensor noise, partial observability, hardware variation, and safety constraints. Those conditions would require separate validation before applying the reported gains to robots or swarm systems.

Integration has four moving parts

Odyssey presents the framework as agnostic to the policy-learning algorithm and world-model architecture. Practical compatibility depends on whether an existing system can expose the data and controls needed by the dual curriculum. An implementation requires:

  • Recorded trajectories containing real states, actions, and outcomes for replay and comparison.
  • A discrepancy metric that measures rollout error and defines the fidelity threshold.
  • A curriculum scheduler that can suspend, reprioritize, and reinstate start states.
  • A retraining budget for targeted world-model updates and repeated validation.

PROWL-2 functions as an orchestration layer around an existing model-based learner, coordinating exploration, policy training, error detection, and repair. Teams evaluating it would need to measure the additional compute and real-environment access against the interactions saved through imagined rollouts.

From training loop to product layer

Odyssey describes its learned simulators as experience machines: environments where agents can collect training experience that would be expensive or unsafe to obtain directly. The company positions PROWL as a fidelity and curriculum layer for foundation world models such as Odyssey-3, with a longer-term goal of training multimodal systems through grounded, open-ended interaction.

Trending
  • No trending articles

Comments

avatar

Next Reads