University of Alberta's Wayfarer Beats DreamerV3 on Atari's Hardest Exploration Games
Wayfarer learns reusable skills from pixels alone, cracking Montezuma's Revenge and Private Eye while beating DreamerV3, Rainbow, and IQN.
- Wayfarer discovers reusable options from raw pixels and beats DreamerV3, Rainbow, and IQN on hard Atari games.
- Uses Laplacian representation learning restricted to features the agent can actually control.
- Three networks (representation, option, control) trained online from a single experience stream.
- Biggest gains on Montezuma's Revenge and Private Eye, both long-horizon exploration games.
- Discovered options generalise to unseen Montezuma rooms without extra training.
- Model-free, single-stream, tabula rasa, no world model or handcrafted features required.
Wayfarer learns long-horizon Atari skills from pixels
University of Alberta researchers have introduced Wayfarer, an online reinforcement-learning agent that discovers temporally extended skills, known as options, from pixel observations. The arXiv preprint reports the strongest aggregate results among the systems evaluated on 10 Atari 2600 games selected for difficult exploration and long-term credit assignment.
Wayfarer learns which parts of an observation the agent can control, extracts large-scale structure from that representation, and converts the resulting features into reusable behaviors. A high-level policy can then choose either a primitive action or a learned option, reducing the number of decisions required to traverse long sequences.
Why flat control stalls
Conventional Atari agents choose among primitive actions such as moving left, moving right, or pressing fire. That approach works well when rewards follow actions quickly. Sparse-reward games create a harder problem because the agent may need to execute hundreds of correct actions before receiving evidence that any earlier decision helped.
An option packages a sequence of actions into a mini-policy with its own initiation and termination conditions. An option might move a character to a ladder and climb down, allowing the control policy to select that complete behavior instead of issuing each directional command separately. This temporal abstraction can accelerate exploration and shorten the chain of decisions between an action and its eventual reward.
Practical option discovery has remained difficult. Earlier methods often relied on small state spaces, handcrafted features, or structured inputs such as emulator RAM. Classical eigenoption research from the same lab used singular value decomposition over compact state features, an approach that does not directly scale to high-dimensional pixel observations.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.