LOOM recipe scales looped Mixture-of-Experts models to 9-12 loops

LOOM recipe scales looped Mixture-of-Experts models to 9-12 loops

Looped Transformers treat recurrent depth as a new scaling axis for language models. By applying the same shared Transformer blocks repeatedly, they increase effective depth without adding parameters. The paper says the benefits remain unclear for large Mixture-of-Experts (MoE) LLMs once compute is matched in FLOPs. Gains from extra iterations shrink quickly and can even turn into degradation, so the additional FLOPs spent on looping buy little. That is why, the authors say, prior work typically settles on two loops.

The paper names two obstacles to scaling looped MoE. The first is that looping inherits and amplifies the curse of depth: hidden-state variance grows with each iteration as residual updates accumulate, which destabilizes deep recurrence and causes representations to drift. The second is expert selection collapse: routers repeatedly select the same experts across loops, so extra iterations add computation without adding computational diversity.

The proposed fix is LOOM, built on one principle: each loop should contribute new computation while keeping the recurrent state stable. For stability, LOOM scales residual updates to bound variance growth and re-injects the input embedding at every loop. For diversity, it uses per-loop routers that engage different experts, plus a Looping Residual that carries earlier outputs forward.

Experiments across models from 100M to 1.7B parameters show stable scaling to 9-12 loops. Under near-iso-FLOP conditions, the 700M model performs best at 5 loops. Against the non-looped baseline, perplexity falls from 18.36 to 16.54, and average zero-shot accuracy rises from 38.84% to 39.53%. Without FLOP matching, the 1.7B model trained on 60B tokens peaks at 9 loops. There perplexity drops from 9.62 to 7.77, and average zero-shot accuracy goes from 42.4% to 47.7%. Code is released on GitHub at github.com/hed-ucas/LOOM.

Key facts

  • LOOM targets two named failure modes of looped MoE: hidden-state variance that grows with each loop, and routers that pick the same experts in every loop.
  • Stability comes from scaling residual updates and re-injecting the input embedding at each loop; diversity comes from per-loop routers and a Looping Residual.
  • Across 100M-1.7B models, the paper reports stable scaling to 9-12 loops, where prior work typically stops at two.
  • At 700M under near-iso-FLOP, 5 loops gives the best result: perplexity 18.36 to 16.54 and average zero-shot accuracy 38.84% to 39.53% versus the non-looped baseline.
  • At 1.7B on 60B tokens, without FLOP matching, the peak is at 9 loops: perplexity 9.62 to 7.77 and accuracy 42.4% to 47.7%.

Why it matters

Looping is an attractive way to get more depth without more parameters, but for MoE models it has mostly stalled at two loops because extra iterations either help little or hurt. This paper offers a diagnosis (variance growth and expert selection collapse) and a recipe aimed at each cause. If the results hold up, recurrent depth becomes a more usable scaling axis for MoE models.

Who it affects

Mainly researchers and engineers working on LLM architecture, particularly MoE and looped or recurrent-depth Transformers. The experiments cover models from 100M to 1.7B parameters, so the direct audience is people training models at those scales or studying how the approach carries to larger ones.

How to use it

The authors released code at https://github.com/hed-ucas/LOOM. The recipe has four pieces: scaled residual updates, input embedding re-injection at every loop, per-loop routers, and a Looping Residual. The reported best loop count depends on the setup: 5 at 700M under near-iso-FLOP, and 9 at 1.7B without FLOP matching.

How solid is it

The numbers come from the paper's abstract, and the reported gains are real but uneven. The result closest to a FLOP-matched comparison (700M, near-iso-FLOP) is modest: perplexity falls by roughly 10% and accuracy rises by about 0.7 points. The larger 1.7B gain (perplexity down roughly 19%, accuracy up about 5.3 points) is explicitly without FLOP matching, so it partly reflects extra compute. The summary does not name the benchmarks behind average zero-shot accuracy.

Risks and caveats

No models larger than 1.7B are reported, so claims about large MoE LLMs remain a problem framing rather than a tested result. The abstract does not say which model size reaches 12 loops or whether 9-12 loops applies to every size. No wall-clock, latency, memory or hardware figures are given, so the practical cost of running many loops is unknown. The training token count for the 700M model and the baseline setups are also not given.

“each loop should contribute new computation while keeping the recurrent state stable”

— LOOM paper abstract, on its guiding principle