Google's Tiny 10K-Parameter Neural Cellular Automata Beats Bigger AI on Maze Reasoning
A Google team shows tiny grids of locally-connected cells can solve Sudoku, mazes, and ARC-AGI through emergent, visible reasoning dynamics.
- Google's Paradigms of Intelligence team shows Neural Cellular Automata can solve mazes, Sudoku, and ARC-AGI through local-only updates.
- A 10K-parameter NCA solves 201x201 mazes at 100% accuracy, beating a 784K-parameter DeepThink baseline using 44x less compute.
- Reaches 60.3% pass@64 on ARC-AGI-1 with a 3.5M parameter model, 63.0% with a 3-model ensemble.
- Trained on 9x9 mazes, generalizes to 500x larger grids with no weight changes by just tiling more cells.
- Niche-Capped Diversity Pruning cuts Sudoku-Extreme inference from 18.4 to 4.6 hours on H100s with no accuracy loss.
- Decentralized architecture enables adaptive compute and self-repair, cutting damage-recovery operations by about half.
Modern reasoning models often use global attention, allowing every token to exchange information with every other token at each layer. A Google research paper tests a more constrained design: Neural Cellular Automata, or NCAs, built from grid cells that communicate only within their immediate 3 × 3 neighborhoods. The system resembles Conway’s Game of Life, except a learned neural network replaces the fixed update rule.
Models with as few as 10,000 parameters match or outperform specialist architectures on maze benchmarks, transfer from 9 × 9 training boards to mazes roughly 500 times larger by area, and reach 60.3% pass@64 on ARC-AGI-1 with 3.5 million parameters. Their intermediate grid states also expose spatial reasoning dynamics as paths, constraints, colors, and shapes propagate across the board.
Reasoning one neighborhood at a time
Full attention over a grid of H × W cells creates pairwise interactions that scale quadratically with the number of positions. It also requires frequent movement of activations between distant units, which raises energy and communication costs. Those requirements fit current accelerators, but they map less naturally onto neuromorphic, distributed, and spatial computing hardware.
An NCA applies the same small multilayer perceptron at every grid location. Parameter count therefore remains fixed as the board grows. For a board of H × W cells running for T iterations, update work scales roughly with HWT, although larger boards usually require more iterations for information to cross them.
- Cell state: Each position stores a learned vector containing visible inputs, hidden features, predictions, and confidence values.
- Local perception: A cell reads its own state and those of its eight immediate neighbors.
- Shared update: The same small network computes a residual change for every active cell.
- Asynchronous execution: A stochastic firing mask updates only a subset of cells during each iteration.
- No absolute coordinates: The model receives no positional embeddings that identify a cell’s global location.
Information can move only one cell per iteration, so distant regions communicate through repeated local updates. This constraint reduces per-step connectivity while creating a latency cost for problems with long-range dependencies.
Local rules produce global behavior
During maze rollouts, activations explore multiple routes in parallel. Subsequent waves propagate backward through dead ends and remove invalid branches until the remaining states encode a solution path. The resulting dynamics resemble distributed pathfinding rather than a serial search procedure explicitly programmed into the architecture.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.