Queen, a 4B chess language model, explains its moves at Grandmaster level

Modern chess engines are silent experts: they play at a superhuman level but offer no explanation for their play. Language models have the opposite problem. They can produce plausible-sounding explanations, but they play chess too weakly for those explanations to be useful. A new paper introduces Queen, a 4B-parameter chess-language model meant to close that gap. The authors say it can explain its moves and plans while playing at the level of a typical Grandmaster.\n\nThe framework has two parts. The first is an encoder-decoder architecture that integrates a silent expert chess encoder with an instruction-tuned language model through cross-attention. It is trained with a question-answering curriculum so that the language model learns to extract chess concepts from the encoder's representations. The second part is an iterative distillation algorithm. Starting from this domain-adapted model, the authors improve its explanations with what they call a natural-language analog of the Bellman update. The model analyzes the positions that arise after its top candidate moves, consolidates those analyses into an explanation of the current position, and that explanation is then distilled back into the model.\n\nOver seven iterations of this loop, the model gains over 900 Elo points, from 1782 to 2697. The authors report that it substantially surpasses all frontier models on both playing strength and puzzle accuracy, despite having three orders of magnitude fewer parameters. They also report that LM-based evaluations find the explanations fluent and approaching GPT-5.6-Sol (high) in coherence.\n\nThe authors close with a broader suggestion: the generality of the architecture and training procedure points to a recipe for applying language models to any domain where silent expert encoders are available, such as games, robotics and computer use.
Key facts
- Queen is a 4B-parameter chess-language model that, per the authors, explains its moves and plans while playing at the level of a typical Grandmaster.
- The design joins a silent expert chess encoder to an instruction-tuned language model via cross-attention, trained with a question-answering curriculum.
- Over seven iterations of a natural-language analog of the Bellman update, with explanations distilled back into the model, its Elo rose by over 900 points, from 1782 to 2697.
- The authors say it substantially surpasses all frontier models on playing strength and puzzle accuracy with three orders of magnitude fewer parameters.
- LM-based evaluations rate the explanations as fluent and close to GPT-5.6-Sol (high) in coherence.
Why it matters
Chess engines are strong but silent, and language models can talk but play poorly. Queen is presented as a way to get both in one small model. The training idea is the novel part: instead of only adding explanations on top of an engine, the model analyzes positions after its own top candidate moves, writes an explanation, and learns from that explanation. The authors report that this loop lifted the model from 1782 to 2697 Elo over seven iterations.
Who it affects
The immediate audience is researchers working on domain-specific reasoning in language models and on explainable game-playing systems. The authors suggest the same recipe could apply to other domains with silent expert encoders, naming games, robotics and computer use. That is a suggestion; the paper reports no experiments in those domains.
How to use it
No release of weights, code or a demo is mentioned, so there is nothing for readers to run yet. The practical takeaway is the method: a cross-attention link between a silent expert chess encoder and an instruction-tuned LM, a question-answering curriculum to teach concept extraction, and a repeated analyse, consolidate and distill loop.
How solid is it
Every figure here is the authors' own report. The headline claims are the 1782 to 2697 Elo climb, surpassing all frontier models on playing strength and puzzle accuracy, and explanations that approach GPT-5.6-Sol (high) in coherence. The explanation quality is judged by LM-based evaluations. The Elo scale or rating pool, the opponents and time controls are not specified, and no numbers are given for puzzle accuracy or coherence. The frontier models used for comparison are not named, apart from GPT-5.6-Sol (high) as the coherence reference.
Risks and caveats
The explanations are not stated to be verified as correct; only fluency and coherence are reported, via LM-based evaluation. A fluent explanation of a chess move is not the same as a faithful one. The base language model and the chess encoder are not named, and no authors or institutions are named. The claim about robotics, games and computer use is a suggestion, not a demonstrated result.
“Modern chess engines are silent experts: they play at a superhuman level, but do not offer explanations for their play.”
— Abstract of the Queen paper