Google's RRSI stops self-improving AI agents memorizing tests

Google's RRSI stops self-improving AI agents memorizing tests

Modern AI agents wrap a fixed language model in a harness: a framework of prompts, workflows, tools, memory and logic that controls what the model sees at each step. The harness decides whether an agent reads the right file before changing it, whether it recovers from a mistake, and whether it delivers its results cleanly. According to a new research paper, much of the recent progress in agents comes from work on the harness, not from new models.

Harnesses used to be tuned by hand: people reviewed failed runs and patched them manually. Newer methods automate the loop by having a language model rewrite the harness again and again, based on feedback from test tasks. The researchers call this a practical form of recursive self-improvement, because the system produces feedback that it uses to optimize the harness, which in turn controls the system's own behavior.

The paper's central finding is that this loop comes with a catch. Because the agent keeps working on the same limited set of test tasks, it ends up memorizing them. Its scores on the training tasks go up, while gains on new, unseen tasks shrink or disappear entirely. The researchers say this happens in several ways: the search memorizes patterns that only fit one particular benchmark, favors candidates that score well purely by chance, and piles on unnecessary complexity that raises the test score without making the agent any better.

The new method, from Google Cloud AI Research and several universities, is called RRSI (Regularized Recursive Self-Improvement of Agent Harnesses). It keeps the harness fully editable but restrains the optimization loop at two points. When the system proposes changes, a budget caps how many independent edits a candidate can bundle at once. That budget shrinks over time: early rounds allow larger rewrites, later rounds only small changes that can be clearly traced to a result. The system also tracks earlier attempts so it does not keep chasing failed ideas, and when progress stalls it deliberately experiments with parts of the harness it has not touched yet.

When deciding which changes become permanent, a critic reviews every proposal and throws out any that hardcode task names, solutions or other benchmark-specific tricks. A further rule accepts higher compute costs only if they come with a measurable performance gain, and components that no longer help are removed.

The team tested RRSI on eight benchmarks spanning coding, agentic office work and engineering design. The underlying model, Claude Opus 4.8, stayed frozen throughout. RRSI was compared with the unmodified baseline harness and four recent optimization methods. According to the paper, RRSI gains up to 14.1 points on the tasks it was trained on and up to 4.7 points on five benchmarks it never saw, with the biggest unseen gain of 4.7 points on JobBench. It also uses about 30 percent fewer tokens at runtime than the unregularized version. Overall performance never fell below the baseline on any of the unseen benchmarks, which typically happens with a harness that has memorized its tasks.

Every method did well on the training tasks, but the results flipped on new ones. Two methods even ended up below the baseline harness. RRSI posted the smallest training gain of all the variants and was the only method that landed well above the baseline on unseen tasks, which is the tradeoff the guardrails are meant to produce. Among the optimized harnesses RRSI needs the fewest tokens and steps, though the unmodified baseline harness is even leaner.

The gains also transferred across models. A coding harness optimized with Gemini 3.5 Flash raised the accuracy of the much weaker Gemini 3.1 Flash Lite from 11.2 to 14.6 points without any modifications, which suggests the mechanisms the system found do not depend on the capability of the model used to discover them.

The authors note that the study only covers harnesses built around frozen models and does not address cases where the model weights change. They conclude that self-improvement only makes AI agents reliably more capable when repeated feedback gets turned into lasting changes. The code is available on GitHub.

The article adds background from other work. Manually designed harnesses often fail to generalize, as tests on ARC-AGI-3 showed: with a purpose-built harness, Opus 4.6 scored 97.1 percent in a familiar environment and 0 percent in an unfamiliar one. Nvidia recently presented SoL-Pi, a related method in which a research agent automatically rebuilds the harness of coding agents, cutting token use by up to 49 percent without a noticeable drop in performance. Shortly before that, Google had agents "dream" about past search runs to improve their search strategy, which also leaves the model unchanged.

Key facts

  • RRSI (Regularized Recursive Self-Improvement of Agent Harnesses) limits overfitting when a language model repeatedly rewrites an agent's harness; the model itself stays frozen.
  • It caps and shrinks the number of independent edits per proposal, and a critic rejects changes that hardcode task names, solutions or other benchmark-specific tricks.
  • Tested on eight benchmarks against the baseline harness and four recent methods: gains of up to 14.1 points on training tasks and up to 4.7 points on five unseen benchmarks.
  • RRSI uses about 30 percent fewer tokens at runtime than the unregularized version, though the unmodified baseline harness is still leaner.
  • A coding harness optimized with Gemini 3.5 Flash lifted Gemini 3.1 Flash Lite from 11.2 to 14.6 points with no changes.

Why it matters

Harness work is where much of the recent progress in agents comes from, according to the paper, and automating it with a model rewriting its own harness is a practical form of recursive self-improvement. This work shows the obvious version of that loop has a flaw: scores on the training tasks climb while gains on new tasks shrink or vanish, because the agent memorizes its tests. RRSI is a way to keep the self-improvement loop while making the gains carry over to tasks the agent has not seen. It also cuts runtime token use compared with the unregularized version, so generalization and compute cost improve together. The authors' framing is that self-improvement makes agents reliably more capable only when repeated feedback becomes lasting changes.

Who it affects

Teams that build agents around a fixed language model and tune the harness (prompts, workflows, tools, memory, logic) are the direct audience, especially anyone automating that tuning with a model in the loop. It also matters to people who read agent benchmark results, since training-set gains from automated optimization may not reflect real capability. The paper's authors are from Google Cloud AI Research and several universities.

How to use it

The code is available on GitHub. In practice the method suggests several design rules for an automated harness optimizer: limit how many independent edits a candidate can bundle, and shrink that limit over time; keep a record of past attempts so the search does not repeat failed ideas, and try untouched parts of the harness when progress stalls; have a critic reject proposals that hardcode task names or solutions; accept higher compute cost only when it buys a measurable gain; and remove components that stop helping. The transfer result is also useful: a coding harness optimized with Gemini 3.5 Flash raised Gemini 3.1 Flash Lite from 11.2 to 14.6 points without modifications, so a harness found with one model may help a weaker one.

How solid is it

The results come from the paper itself, as reported by The Decoder. Testing covered eight benchmarks in coding, agentic office work and engineering design, with comparisons against the unmodified baseline harness and four recent optimization methods. Performance never fell below the baseline on any unseen benchmark, while two of the other methods did end up below it. The strongest unseen gain was 4.7 points on JobBench. Note that RRSI had the smallest training gain of all the variants; the guardrails are meant to produce exactly that tradeoff. The source does not specify whether the 14.1 and 4.7 point figures are percentage points or relative gains, and the 30 percent token saving is measured against the unregularized version, not the baseline harness.

Risks and caveats

The authors say the study only covers harnesses built around frozen models and does not address cases where the model weights change. The reported unseen gains are modest, up to 4.7 points on five unseen benchmarks. The optimized RRSI harness is the leanest among the optimized variants, but the unmodified baseline harness uses fewer tokens and steps still, so the efficiency claim holds only against other optimized methods. The source does not name the four competing methods or list the eight benchmarks beyond JobBench, which limits how far a reader can judge the comparison. ARC-AGI-3, Nvidia's SoL-Pi and Google's "dream" work are background from other studies, not RRSI results.