DISCO splits long context across worker LLMs to curb context rot

DISCO splits long context across worker LLMs to curb context rot

The paper starts from a familiar complaint. Large language models advertise million-token context windows, yet reasoning quality often collapses as inputs grow. The authors call this context rot. They blame a structural entanglement in monolithic architectures: the search burden of finding the relevant material in a huge context (contextual grounding) uses up the representational capacity the model needs for complex reasoning.

Their answer is DISCO, short for Grounding-Reasoning Disaggregation via DIStributed long COntext scaling. The design is inspired by distributed computing frameworks such as Apache Spark. The long context is partitioned across a fleet of Worker LLMs, which do nothing but parallel, localized grounding over their own slice. A central Driver LLM orchestrates the run. It is trained with reinforcement learning (GRPO) to optimize planning, and it works in two steps: it dynamically maps the query into atomic extraction tasks for the workers, then reduces the gathered evidence to synthesize a final answer. The map-and-reduce vocabulary follows the Spark analogy.

The authors' central claim is that isolating reasoning from raw context noise effectively eliminates context rot. The reported results are three. On RULER-QA at 1M tokens, DISCO maintains 78.4% accuracy where standard baselines collapse. On LongBench v2 it outperforms full-context models by up to 9.8 points. And it matches frontier models such as Gemini-3-Pro-Preview while reducing inference costs by over 80%. The authors conclude that this establishes a highly efficient paradigm for robust long-context inference.

The source is a short abstract. It does not name the baselines, the model sizes or the number of workers, and it does not say what the 80% cost reduction is measured against.

Key facts

  • DISCO (Grounding-Reasoning Disaggregation via DIStributed long COntext scaling) splits a long context across a fleet of Worker LLMs that only do parallel, localized grounding.
  • A central Driver LLM, trained with reinforcement learning (GRPO), maps the query into atomic extraction tasks and then reduces the gathered evidence into a final answer.
  • On RULER-QA at 1M tokens, DISCO reportedly keeps 78.4% accuracy where standard baselines collapse.
  • On LongBench v2 it beats full-context models by up to 9.8 points, and the authors say it matches frontier models like Gemini-3-Pro-Preview.
  • The authors report inference costs cut by over 80%; the baseline for that figure is not stated.

Why it matters

Million-token context windows are a headline feature, but the authors argue that reasoning quality often collapses as inputs grow, which they call context rot. DISCO tackles this by changing the architecture rather than the model: finding evidence and reasoning over it are handed to different components. The authors say this separation effectively eliminates context rot, and they describe the result as a highly efficient paradigm for long-context inference. The idea borrows from distributed computing frameworks like Apache Spark, with workers doing the parallel scanning and a driver doing the planning and final synthesis.

Who it affects

The work speaks to anyone who needs LLMs to answer questions over very long inputs, where a single model has to both search the text and reason about it. It also concerns people who pay for long-context inference, since the authors report a cost reduction of over 80%. The abstract does not say who would deploy the system or in what products.

How to use it

The abstract describes the method but not a way to run it. Conceptually, a user would split the long input across Worker LLMs, let a Driver LLM turn the question into small extraction tasks, and have it combine the returned evidence into an answer. The abstract does not mention code, a model release or a dataset release, and it does not give the sizes or number of the Worker LLMs or the Driver's base model.

How solid is it

All numbers come from the authors' own summary of their results, and only the abstract is available here. The headline figure is 78.4% accuracy on RULER-QA at 1M tokens, versus baselines that the authors say collapse. The abstract does not say which baselines these are or what accuracy they reach. The 9.8-point margin on LongBench v2 is an upper bound ("up to") over full-context models, and it is a gap in points, not a percentage. The abstract does not say which model or setting produces it. The claim of matching Gemini-3-Pro-Preview is not tied to a named benchmark.

Risks and caveats

The over-80% cost reduction has no stated baseline, so it cannot be read as a saving against Gemini-3-Pro-Preview or against full-context models specifically. The claim that DISCO effectively eliminates context rot is the authors' own wording and rests on the benchmarks they report. No authors, institutions or dates are given in the source, and no independent replication is mentioned.