Recurrent Looped Transformer generalizes parity to 256 bits

Recurrent Looped Transformer generalizes parity to 256 bits

The paper starts from a limitation of standard Transformers. State tracking requires an update at every input, but the depth a Transformer applies to each token is fixed regardless of sequence length. The authors introduce the Recurrent Looped Transformer (RLT), which splits its layers between a parallel causal encoder and a recurrent decoder. At each token, the decoder merges the encoder output with the previous token's final decoder state. The computation path therefore grows with sequence length, while the per-token cost stays fixed.

The evaluation covers six algorithmic tasks. The authors compare five splits of eight layers (RLT variants) against an eight-layer Transformer, over three seeds.

On parity, models were trained on at most 40 bits. Two RLT splits generalize to 256 bits with 100% accuracy in every seed, while the Transformer stays at chance. On swap-based S_5 permutation tracking at eight times the training length, RLT reaches 97% final-state accuracy versus under 1% for the Transformer, and accuracy increases with decoder depth. On modular arithmetic beyond the training lengths, RLT reaches up to 93% versus 33% for the Transformer.

Ablations show the gains depend on the feedback: removing it drops parity and swap-based S_5 to chance at every split.

The paper also tests a cheaper variant that updates the feedback once per four-token chunk, which lets known tokens in a chunk run in parallel. That keeps 64-bit parity at 99%. Permutation tracking, however, depends on per-token feedback: chunking lowers length-64 swap-based S_5 accuracy from 100% to 20%.

Key facts

  • RLT splits its layers between a parallel causal encoder and a recurrent decoder that merges the encoder output with the previous token's final decoder state at each token.
  • Trained on at most 40 bits, two RLT splits reach 100% parity accuracy at 256 bits in every seed; the Transformer stays at chance.
  • On swap-based S_5 permutation tracking at eight times the training length, RLT hits 97% final-state accuracy versus under 1% for the Transformer.
  • On modular arithmetic beyond training lengths, RLT reaches up to 93% versus 33% for the Transformer.
  • Removing the feedback drops parity and swap-based S_5 to chance at every split; four-token chunking keeps 64-bit parity at 99% but cuts length-64 S_5 from 100% to 20%.

Why it matters

The paper targets a structural limit of Transformers: the depth applied to each token is fixed regardless of sequence length, while state tracking needs an update at every input. RLT addresses this by passing the previous token's final decoder state forward, so the computation path grows with sequence length at a fixed per-token cost. The reported results are large on the chosen tasks, such as 97% versus under 1% on swap-based S_5 at eight times the training length.

Who it affects

The work is aimed at researchers studying length generalization and state tracking in sequence models. All six tasks are algorithmic, and the source mentions no results on natural-language or real-world tasks, so the practical reach beyond these benchmarks is not shown.

How to use it

Nothing here is a product or a ready tool. The practical takeaways are design choices from the ablations. Per-token feedback is needed for permutation tracking, while updating feedback once per four-token chunk lets known tokens in a chunk run in parallel and still keeps 64-bit parity at 99%. On swap-based S_5, deeper decoders gave higher accuracy. No code release or open weights are mentioned.

How solid is it

This is a paper abstract with no outside validation. The setup is controlled: five splits of eight layers are compared with an eight-layer Transformer over three seeds, and the 100% parity result holds in every seed. Ablations back the claim that the feedback drives the gains. The abstract names no authors or institutions. Which two of the five splits reached 100% parity is not stated, and the 93% and 33% modular arithmetic figures do not state the sequence length tested. The 97% S_5 result is given without a per-seed breakdown or variance.

Risks and caveats

The evidence covers only algorithmic tasks, and all comparisons are against a single eight-layer Transformer baseline. The best results depend on per-token feedback, which gives up the full parallelism of chunked updates: chunking lowers length-64 swap-based S_5 from 100% to 20%. No model sizes beyond the eight-layer budget, parameter counts or training compute are stated.