LoopOPD lets looped language models learn from their own deeper loops

Looped Language Models (LoopLMs) scale reasoning by reusing the same shared parameters across several recurrent computation steps, which makes them parameter efficient. The paper says that post-training them well is still hard. Existing approaches either rely on reward-based supervision, which is sparse or costly to extend across loops, or depend on external teachers or privileged information. That leads to limited teacher availability or a mismatch between the teacher's context and the student's.

The authors introduce LoopOPD, a cross-loop on-policy distillation framework. Its idea is to treat the extra recurrent computation inside a LoopLM as its own source of supervision. A frozen terminal-loop policy acts as a "compute privileged" teacher for an intermediate-loop student. The student generates the rollouts, and the teacher supplies dense supervision on them, so no external teacher and no privileged information are needed.

The paper also proposes Dynamic LoopOPD (D-LoopOPD). Instead of keeping the terminal-loop teacher frozen, it continually refreshes that teacher as the shared model parameters are updated. The authors say this enables recurrent self-improvement.

On the theory side, the authors characterize how distillation updates propagate across loop depths. They derive sufficient conditions under which a single update gives simultaneous local improvement at both loop depths.

In experiments on Ouro-Thinking models, LoopOPD improves mathematical reasoning, and D-LoopOPD yields further gains through the dynamic teacher updates. The models were trained only on mathematical data, yet the paper reports that they also improve on general reasoning and code generation benchmarks. The authors conclude that recurrent computation can serve as an effective source of supervision for LoopLMs. They say code and model checkpoints will be released upon acceptance.

Key facts

  • LoopOPD uses a frozen terminal-loop policy as a teacher for an intermediate-loop student, on rollouts the student generates itself.
  • The method needs no external teacher and no privileged information; the extra recurrent computation inside the model supplies the supervision.
  • D-LoopOPD refreshes the terminal-loop teacher as the shared parameters change, and the paper reports further gains over LoopOPD in experiments on Ouro-Thinking models.
  • Training used only mathematical data, but the paper reports improvements on general reasoning and code generation benchmarks too.
  • The authors derive sufficient conditions for a single update to improve both loop depths locally; code and checkpoints are promised upon acceptance.

Why it matters

Looped models get more reasoning out of a fixed set of parameters by running them for more steps, but they are hard to post-train. This paper offers a way to do it without outside help: the model's deeper loops teach its shallower ones. If the approach holds up, a looped model can improve using only its own extra computation as the training signal, which the authors call recurrent self-improvement.

Who it affects

The work is aimed at researchers building and post-training looped language models, and the experiments use Ouro-Thinking models. Teams that find external teachers or privileged information hard to obtain for this kind of architecture are the natural audience for a method that needs neither.

How to use it

There is nothing to run yet. The authors say their code and model checkpoints will be released upon acceptance. Until then, the abstract describes the recipe only in outline: distill a frozen terminal-loop policy into an intermediate-loop student on the student's own rollouts, and in the dynamic variant refresh the teacher as the shared parameters are updated.

How solid is it

This is an arXiv paper, and the account here rests on its abstract. The abstract gives the claims and a theoretical result in the form of sufficient conditions for simultaneous local improvement at both loop depths. It reports no numbers: no accuracy figures, benchmark names, baselines or sizes of gains. The claim of gains beyond math, on general reasoning and code generation, is therefore stated but not quantified here.

Risks and caveats

The theoretical conditions are sufficient, not necessary, and they guarantee local improvement, not improvement in general. The abstract does not state the sizes or loop counts of the Ouro-Thinking models, or which loop depths serve as student and teacher in the experiments. The released code and checkpoints are conditional on acceptance, so independent reproduction has to wait.

“demonstrating that recurrent computation can serve as an effective source of supervision for LoopLMs.”

— Paper abstract, arXiv 2610.10623