VISTA Gives Claude Opus 5.0 Visual Memory and Hits Perfect ARC-AGI-3

VISTA gives multimodal models a lossless visual memory they can zoom into, lifting Claude Opus 5.0 to a perfect score on ARC-AGI-3.

·
·
·
VISTA Gives Claude Opus 5.0 Visual Memory and Hits Perfect ARC-AGI-3PRO
  • VISTA is a minimal visual harness that lifts Claude Opus 5.0 to a perfect 100 RHAE on ARC-AGI-3.
  • The system solves all 25 public games using 57.4% fewer actions than first-time human players.
  • Core idea: raw PNG observations plus a lossless visual memory the model can zoom into on demand.
  • Two scratch files, GUIDE.md and WORKING.md, let the agent maintain a revisable model of each game.
  • Also beats minimal-harness baselines on GameWorld, AI GameStore, and BabyVision with no task-specific tuning.
  • Code and setup instructions available at github.com/joshhhhhan/VISTA.

VISTA gives visual agents a memory, then clears ARC-AGI-3

The VISTA paper reports a perfect 100.00 RHAE score across all 25 public ARC-AGI-3 games by placing a lightweight visual harness around Claude Opus 5.0. The system uses raw screenshots, on-demand access to prior frames, and editable notes, without program synthesis or game-specific logic. Experiments also pair the harness with GPT-5.6 and an open-weight GLM model.

ARC-AGI-3 asks an agent to infer each game’s rules and objective through interaction. A completed level earns full credit when the agent stays within the first-time-human action baseline, while less efficient completions receive partial credit and failures receive zero. A 100.00 aggregate score means the agent completed every public game within the full-credit action budget. For agent developers, the result shows how much observation format, memory retrieval, and workspace design can affect performance while model weights remain fixed.

Pixels stay pixels

VISTA preserves the model’s strongest input format throughout the task. Current game states arrive as PNG images, prior frames remain available at full resolution, and the model reasons in free-form language instead of receiving a flattened text grid.

  • Native visual input: The model sees rendered game frames rather than serialized cells.
  • Lossless visual memory: Every returned frame is stored in its original form and indexed by turn and frame number.
  • Selective retrieval: The model requests specific historical frames or regions when earlier evidence becomes relevant.

Older images stay outside the active prompt until requested, which preserves exact visual evidence without filling the context window with every frame from the episode.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads