Princeton's Queen Teaches a Grandmaster Chess Engine to Explain Its Moves
Princeton's Queen couples a Leela chess network to SmolLM3 via cross-attention, then self-improves through an AlphaZero-style natural language distillation loop.
- Princeton released Queen, a 4B chess-language model that plays at grandmaster level and explains moves.
- Architecture bolts a Leela Chess Zero encoder to SmolLM3 via Flamingo-style cross-attention blocks.
- Bridge parameters learn via a QA curriculum covering static and dynamic board features.
- Self-improvement uses an AlphaZero-style Bellman update expressed in natural language over search trees.
- Seven iterations lift Elo from 1782 to 2697, beating frontier LMs 1000x larger.
- Code, paper and checkpoints available on Hugging Face and arXiv.
Queen gives a grandmaster-level chess engine a voice
Princeton researchers have introduced Queen, a roughly 4-billion-parameter model that pairs grandmaster-level chess play with plain-English move analysis. Across seven self-improvement rounds, its reported Lichess blitz rating rose from 1782 to 2697, and it outperformed every frontier language model tested on both playing strength and puzzle accuracy.
Strong chess engines such as Stockfish and Leela Chess Zero produce moves and numerical evaluations with little natural-language guidance. General-purpose language models can generate fluent analysis, yet historically have struggled with legal moves, board state, and tactics. Queen, short for Quality Explanation and Evaluation Network, connects an expert chess encoder to a language decoder so that both systems contribute to each response.
The narrow bridge to SmolLM3

Queen combines an Lc0 BT5 board encoder with a fine-tuned SmolLM3 decoder. The encoder converts a chess position into internal representations of the board; the decoder uses those representations to generate moves and prose analysis token by token.
The connection follows DeepMind’s Flamingo architecture. Embeddings from layer k of the Lc0 network enter a Flamingo cross-attention block before layer 2k of the language model. During domain adaptation, the previously trained components remain frozen and only the cross-attention bridge is updated. This setup preserves the chess network’s learned position evaluation while teaching the language model where to retrieve board information.
A four-stage question-and-answer curriculum trains the bridge with progressively harder tasks:
- Static current: identify pieces, squares, attacks, and other facts about the present position.
- Dynamic current: analyze legal moves and tactical interactions, including which pieces can give check.
- Static future: reconstruct the board after a supplied move sequence.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.