UCL's Memento 3 Beats Human Efficiency by 56% on Every ARC-AGI-3 Puzzle

A frozen LLM agent learns by rewriting its own natural-language rulebook and compiling it into verified code, clearing every public ARC-AGI-3 game.

·
·
·
UCL's Memento 3 Beats Human Efficiency by 56% on Every ARC-AGI-3 PuzzlePRO
  • Memento 3 reaches the ARC-AGI-3 RHAE ceiling of 100.0 across all 25 public games with a frozen LLM.
  • Agent uses 7,518 actions versus the 17,135-action human baseline, roughly 44 percent efficiency.
  • World model is a Markdown rulebook compiled into Python, revised through observe-reflect-revise-compile-verify loops.
  • Revisions accepted only when LLM judges faithfulness and cell-exact replay reproduces observed transitions.
  • Atari Pong controller derived from the rulebook wins 21:0 in three episodes with no LLM calls at inference.
  • Paper available on arXiv; prior Memento code on GitHub.

Memento 3 compiles an LLM’s game rules into executable code

Researchers from University College London and Huawei’s Noah’s Ark Lab describe Memento 3, an agent that builds and revises an external world model while keeping its large language model frozen. The agent records its current understanding of a game in Markdown, compiles those rules into Python, and checks the resulting program against observed state transitions. When a prediction fails, it updates the rulebook and recompiles the engine without changing the LLM’s weights.

On the 25 public ARC-AGI-3 games, the authors report that Memento 3 cleared every level with a mean Relative Human Action Efficiency score of 100.0, the metric’s ceiling. It used 7,518 actions, or 44% of the 17,135-action human baseline. RHAE measures action efficiency against a human reference and caps scores at 100, so the metric does not distinguish between matching and exceeding that reference.

Why plausible models fail

ARC-AGI-3 is an interactive, game-based successor to the ARC visual reasoning puzzles. An agent must infer each game’s objects, controls, dynamics, and goals from observations gathered while playing. Effective planning therefore depends on a world model that predicts how the environment will respond to an action.

The paper identifies identifiability as the central problem. A limited interaction history can support several explanations that reproduce every observed transition while making different predictions about unseen states. An agent may appear accurate until a new level exposes an incorrect assumption about movement, geometry, or control behavior.

Common LLM-agent designs retain trajectories in the context window or update model weights from experience. Memento 3 stores its evolving hypotheses in persistent external artifacts, where later observations can revise or reject them. This gives the agent a versioned model that survives context limits and can be inspected independently of the LLM.

A rulebook the agent can execute

The authors call the design Code as Model. Three artifacts separate evidence, interpretation, and execution:

Artifact Contents Purpose
Interaction history Actions, observations, and resulting transitions

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads