Memento 3 paper: frozen-LLM agent clears all 25 public ARC-AGI-3 games

Memento 3 is a paper about getting a frozen LLM agent to learn how an unfamiliar environment works. The authors start from a familiar problem: limited observations can support several world models that explain past interactions equally well but predict different outcomes in states the agent has not yet seen. Memento 3 builds on the earlier Memento series and gives the agent an external memory so it can keep learning explicit world models while the underlying model never changes.
The design has three parts. First, the agent maintains a natural-language rulebook as persistent semantic memory. It records revisable hypotheses about how the environment behaves and deliberately leaves unknown aspects underspecified. Second, it compiles that rulebook into executable code, which it then uses for prediction and planning. Third, a continual loop of observation, reflection, rule revision, compilation and verification uses prediction errors to refine both the rulebook and the code.
There is a gate on every update. New code is accepted only when two conditions both hold: the LLM judges it faithful to the rulebook, and cell-exact replay reproduces the observed transitions.
The authors frame this as an investigation of a model-based route to recursive self-improvement (RSI). In their description, the agent explores the environment on its own, revises its world model, and uses verified updates to guide later interaction and learning, while the underlying LLM stays fixed. They also describe a population extension that keeps several world models in parallel, shares interaction evidence between them and uses their predictions to guide exploration.
The reported results come from two settings. On ARC-AGI-3, the single-model agent clears every level of all 25 public games, achieves a mean Relative Human Action Efficiency (RHAE) of 100.0 and uses 44% of the human action count. In an Atari Pong case study, a learned feedback controller wins 21:0 in each of three evaluated episodes with different openings, without further LLM calls.
Key facts
- Memento 3 lets a frozen LLM agent keep a natural-language rulebook as persistent memory and compile it into executable code for prediction and planning.
- Prediction errors drive a loop of observation, reflection, rule revision, compilation and verification; updated code is accepted only if the LLM judges it faithful to the rulebook and cell-exact replay reproduces the observed transitions.
- On ARC-AGI-3 the single-model agent clears every level of all 25 public games, with a mean RHAE of 100.0 and 44% of the human action count.
- In an Atari Pong case study, a learned feedback controller wins 21:0 in each of three evaluated episodes, without further LLM calls.
- A population extension runs multiple world models in parallel and shares interaction evidence; the abstract describes its design but gives no results for it.
Why it matters
The paper tackles a hard question for agents: how to learn what an environment does when observations are limited and several explanations fit the data. Its answer keeps the LLM fixed and puts the learning in an external, inspectable rulebook that is compiled into code and checked against what actually happened. The authors present this as a model-based route to recursive self-improvement, and the claim is framed as something they investigate. The headline number is a full clear of all 25 public ARC-AGI-3 games.
Who it affects
Mainly researchers working on agents, world models and memory for LLM systems, and anyone following ARC-AGI-3 as a benchmark of learning in unfamiliar environments. The approach keeps the underlying LLM unchanged, so it is of interest to teams who cannot or do not want to retrain a model.
How to use it
This is a research preprint. The source describes the method and reports results but mentions no code, model or dataset release, so there is nothing to run yet. For practitioners the transferable idea is the structure: write down revisable hypotheses in plain language, compile them into code, test that code by replaying observed transitions, and accept changes only when a check passes.
How solid is it
All figures come from the paper's own abstract, and the authors report them themselves. The abstract gives no baselines or competing systems for ARC-AGI-3, so it states no comparison with other agents. The 25 games are the public ones, and the source does not say whether results on any private or held-out set exist. The Pong result covers three episodes in one case study, and no wider Atari results are reported. The abstract does not name the underlying LLM, and it gives no compute cost, number of LLM calls or wall-clock time for the ARC-AGI-3 runs.
Risks and caveats
The source does not say that recursive self-improvement was achieved; it says the process is investigated as a route to it, with the LLM held fixed. The population extension is described by design only, with no results. The 44% figure is a share of the human action count, and the source does not say whether it is a mean, median or total. It also does not explain the scale of RHAE beyond the reported mean of 100.0. Treat the Pong win as a case study rather than a benchmark result.
“Updated code is accepted only when the LLM judges it faithful to the rulebook and cell-exact replay reproduces the observed transitions.”
— Memento 3 paper abstract