REMORY adds soft memory tokens to LLM context compaction

REMORY adds soft memory tokens to LLM context compaction

Long-horizon agents compact their history so they can keep working within a finite context window. The authors of a new paper argue that a textual summary alone may not support every subsequent decision the agent has to make.

Their answer is REMORY, a neural memory network that supplements the summary with a bounded sequence of soft memory tokens. Given the full history and the summary, the network learns to generate tokens that help a frozen LLM approximate the continuation it would produce if it could see the full history. The tokens are conditioned on the summary and appended after it, which the paper describes as an analogue of a residual connection along the sequence dimension.

The abstract reports results on two kinds of tests. On SummHay, REMORY improves source attribution at nearly unchanged insight coverage, and it approaches the full-context joint score while using only 5.2% of the input positions. Across long-horizon agent benchmarks, Qwen3.8-27B and GLM-5.3-Flash show consistent gains with residual memory. Both models also show substantially fewer repeated tool outputs and tool errors on BrowseComp and Terminal-Bench 2.1.

The abstract gives no absolute scores for any benchmark, so the size of the gains is not quantified.

Key facts

  • REMORY is a neural memory network that adds a bounded sequence of soft memory tokens after an agent's textual summary.
  • The network is trained so that a frozen LLM approximates the continuation it would produce with the full history.
  • On SummHay, REMORY improves source attribution at nearly unchanged insight coverage and approaches the full-context joint score using only 5.2% of the input positions.
  • Qwen3.8-27B and GLM-5.3-Flash show consistent gains on long-horizon agent benchmarks with residual memory.
  • Both models make substantially fewer repeated tool outputs and tool errors on BrowseComp and Terminal-Bench 2.1.

Why it matters

Agents that run for a long time have to shrink their history to fit the context window, and the usual tool is a text summary. The paper's premise is that a summary alone may not carry everything the agent needs for later decisions. REMORY keeps the summary and adds a small learned layer of soft tokens on top of it, placed after the summary the way a residual connection adds to a layer's output. The target is concrete: make a frozen LLM behave as it would with the full history, at a fraction of the input length.

Who it affects

The work is aimed at people building long-horizon agents that compact their history, such as browsing and terminal agents. The reported tests use Qwen3.8-27B and GLM-5.3-Flash, so the evidence concerns those two models. Because the LLM stays frozen, the approach is framed as something added around the model rather than a change to the model itself.

How to use it

Nothing in the abstract is directly usable yet: no code, model release or licence is mentioned. The method as described is a pipeline step. Compact the history into a text summary as before, run the memory network over the history and summary to produce soft memory tokens, and append those tokens after the summary before the frozen LLM continues.

How solid is it

This is a preprint, and everything above comes from its abstract, so the numbers are the authors' own. The abstract names no authors or institutions. Only one figure is given: 5.2% of the input positions on SummHay. Gains on the agent benchmarks are described as consistent and the drop in tool errors as substantial, but neither is quantified, and no absolute scores are given for SummHay, BrowseComp or Terminal-Bench 2.1.

Risks and caveats

The headline result is that REMORY approaches the full-context joint score on SummHay, not that it matches it. The number of soft tokens, how the network is trained and what that costs are not stated. No comparison with other context-compaction or memory methods is mentioned, so it is unclear how much of the benefit is specific to this design. Results cover two models and a handful of benchmarks.

“The tokens are conditioned on the summary and appended after it, forming an analogue of a residual connection along the sequence dimension.”

— REMORY paper abstract