reViT paper claims one recurrent block matches a full-depth vision transformer

The paper introduces reViT, a vision Transformer built around one block that is applied recurrently instead of stacking many distinct blocks. The authors' central claim is that this single recurrent block can match the accuracy of a full-depth vision encoder at comparable inference FLOPs, without intermediate feature distillation.
The obvious weakness of reusing one block is that every depth would apply the same transformation. reViT addresses this by making the feed-forward network (FFN) at each recurrent depth a convex combination of a small shared expert bank. A continuous, normalized-depth coordinate programs that mixture, which defines a resampleable trajectory through FFN parameter space. In other words, the depth the block is currently at decides which blend of the shared experts it uses.
The design is evaluated in two regimes: supervised ImageNet-1k training and distillation from a DINOv2 teacher. Across both, controlled adaptations identify weight-space merging as the strongest tested mixture-of-experts family at a matching one-FFN budget, ahead of the token-dispatch and output-mixture alternatives.
Two headline results are reported. Trained from scratch, reViT-B/16 attains DeiT III accuracy with about 70% fewer stored parameters. And an 8-experts model distilled using only the teacher's output features retains nearly all of its DINOv2 teacher's linear-probe accuracy and transfers across classification, segmentation, and depth prediction.
Two practical features follow from the design. Elastic-depth training lets one checkpoint operate at multiple tested depths by resampling the same normalized coordinate interval. For fixed-depth deployment, the recurrent block can be materialized as a conventional dense graph, which removes online routing and merging without changing the one-FFN-per-depth compute, but expands deployment storage.
Key facts
- reViT applies a single Transformer block recurrently and gives each depth its own FFN as a convex combination of a small shared expert bank, programmed by a continuous normalized-depth coordinate.
- Trained from scratch, reViT-B/16 attains DeiT III accuracy with about 70% fewer stored parameters.
- An 8-experts model distilled from a DINOv2 teacher using only its output features retains nearly all of the teacher's linear-probe accuracy and transfers across classification, segmentation, and depth prediction.
- Weight-space merging was the strongest tested MoE family at a matching one-FFN budget, ahead of token-dispatch and output-mixture alternatives.
- One elastic-depth checkpoint can run at multiple tested depths; for fixed-depth deployment the block can be materialized as a dense graph at the cost of more storage.
Why it matters
Most vision encoders get their capacity from a stack of distinct blocks. This work argues that one block, reused across depth, can do the same job at comparable inference FLOPs, and that the missing depth-specific behaviour can be restored cheaply with a small shared bank of FFN experts. If the result holds up, the saving shows up in stored parameters: about 70% fewer for reViT-B/16 than DeiT III at the same accuracy. The paper also reports that the approach works both when training from scratch and when distilling from a DINOv2 teacher.
Who it affects
Mainly researchers working on vision transformer design, parameter-efficient architectures and distillation. The reported gain is in stored parameters, so teams that care about checkpoint size are the natural audience. The elastic-depth feature, one checkpoint running at several depths, is relevant to anyone who wants to trade depth for cost at deployment time.
How to use it
This is a research result, not a product. No code, weights or release are mentioned. Conceptually, the recipe is: reuse one Transformer block recurrently, represent the FFN at each depth as a convex combination of a small shared expert bank, and drive the mixture with a normalized-depth coordinate. For fixed-depth deployment the block can be materialized as a conventional dense graph, which removes online routing and merging and keeps the one-FFN-per-depth compute, but expands deployment storage.
How solid is it
All claims come from the authors' own experiments as summarized in the paper's abstract, and the entry on Hugging Face Papers is a preprint listing. The evaluation covers two regimes, supervised ImageNet-1k training and DINOv2 distillation, and the comparison among MoE families was run as controlled adaptations. The abstract names no authors or institutions. No absolute accuracy figures (ImageNet top-1, linear-probe, segmentation or depth metrics) are given, and neither are FLOP or parameter counts; only the relative figure of about 70% fewer stored parameters versus DeiT III is stated.
Risks and caveats
The 70% parameter saving applies to the from-scratch reViT-B/16 versus DeiT III, not to the distilled 8-experts model and not to compute. Retaining "nearly all" of the DINOv2 teacher's linear-probe accuracy is not quantified. Weight-space merging is the strongest only among the MoE families the authors tested, at a matching one-FFN budget. The specific depths at which the elastic-depth checkpoint was tested are not listed. Materializing the block as a dense graph for fixed-depth deployment expands storage, and the size of that increase is not quantified.
“a single Transformer block, applied recurrently, can match the accuracy of a full-depth vision encoder at comparable inference FLOPs without intermediate feature distillation”
— reViT paper abstract (Hugging Face Papers 2610.12448)