Context Language Models let AI models manage their own context as a file

A paper on Hugging Face introduces Context Language Models (CLMs): language models that natively manage their own context. The mechanism is simple to state. The context is treated as a file, and the model is allowed to make unrestricted updates to that file. According to the authors, this lets the model learn what is most important to keep in context, and it extends naturally to multi-agent systems, where several agent contexts coexist as files.
The first set of results concerns CLMs built zero-shot on top of existing models. The authors say these outperform state-of-the-art (SOTA) context management strategies across a variety of tasks. On BrowseComp-Plus they report 11.4% higher accuracy with 21.5% fewer FLOPs. On 12-hour EdgeBench they report 5% higher scores with 59% fewer FLOPs. On a 24-hour multi-repository agent-swarm task they report 65% greater improvement with the same compute.
The second argument is about learning. Because context management moves from external harness control to intrinsic model behaviour, CLMs enable both in-context and parametric learning of context-management strategies. For the in-context route, the authors show that CLMs can be steered with natural-language instructions evolved through a standard skill-optimization loop. That improved held-out accuracy by up to 35.9 points on a context-management task while reducing compute. For the parametric route, they introduce an online reinforcement learning method for CLMs, which improved Qwen3.5-9B performance on BrowseComp-Plus by 47.6% while using 12% fewer FLOPs.
Finally, the authors co-design a serving optimisation called Suffix Cache Reuse for CLM serving. It further reduces server-side compute by 35% relative to standard SGLang at matched performance.
Key facts
- Context Language Models (CLMs) treat the context as a file that the model can update without restriction, rather than leaving context management to an external harness.
- Zero-shot CLMs built on existing models are reported to beat SOTA context management: 11.4% higher accuracy with 21.5% fewer FLOPs on BrowseComp-Plus, and 5% higher scores with 59% fewer FLOPs on 12-hour EdgeBench.
- On a 24-hour multi-repository agent-swarm task, zero-shot CLMs show 65% greater improvement with the same compute.
- Steering CLMs with natural-language instructions evolved through a skill-optimization loop improved held-out accuracy by up to 35.9 points on a context-management task; online RL lifted Qwen3.5-9B on BrowseComp-Plus by 47.6% with 12% fewer FLOPs.
- Suffix Cache Reuse, a serving co-design, cuts server-side compute by a further 35% relative to standard SGLang at matched performance.
Why it matters
Long-running agents have to decide what to keep in their context and what to drop. Today that decision usually sits in an external harness. This paper moves it into the model itself, so the model learns what is most important to maintain. The reported payoff is both better results and lower compute: for example, 11.4% higher accuracy with 21.5% fewer FLOPs on BrowseComp-Plus, and 59% fewer FLOPs on 12-hour EdgeBench. The same design is said to extend to multi-agent systems, where each agent's context is a file.
Who it affects
The work is aimed at people building agentic systems that run for a long time, such as the 12-hour and 24-hour tasks in the benchmarks, and at teams that currently rely on harness-side context management strategies. The serving result, a 35% saving in server-side compute against standard SGLang at matched performance, is relevant to anyone running inference for such agents.
How to use it
The abstract describes building CLMs zero-shot with existing models, so no training is needed for the first set of results. Two routes to improve behaviour are described: steering with natural-language instructions evolved through a standard skill-optimization loop, and an online reinforcement learning method for CLMs. Suffix Cache Reuse is the proposed serving optimisation. No code or model release is mentioned.
How solid is it
The numbers come from the paper's own abstract and are reported by the authors, who compare against unnamed SOTA context management strategies. The abstract does not say whether the 11.4%, 5% and 47.6% gains are percentage points or relative improvements. It also does not name the SOTA baselines or say which base models were used for the zero-shot CLMs on BrowseComp-Plus, EdgeBench or the agent-swarm task. The 35.9-point figure is a maximum ("up to") on one context-management task.
Risks and caveats
The abstract names no authors or institutions, and the headline figures cannot be checked against baselines it does not describe. It does not describe the file format or how the model's file updates work in detail. It does not say what FLOPs baseline the percentages refer to beyond "SOTA context management strategies", nor whether the 35% compute saving from Suffix Cache Reuse applies to all workloads. Letting the model make unrestricted updates to its own context is the core design choice, and the abstract gives no detail beyond the reported benchmark results.
“We implement this by treating the context as a file and allowing the model to make unrestricted updates to this file.”
— Paper abstract, Context Language Models