TT-VidT video pretraining leads four benchmarks with fewer FLOPs

TT-VidT video pretraining leads four benchmarks with fewer FLOPs

The paper starts from a complaint about how video self-supervised learning is compared. Papers often evaluate complete training recipes rather than isolating the method itself, so architecture, objective, data exposure, schedule, scale and decoder capacity can all vary at once. That makes it hard to tell which choices produce motion-prioritized representations, meaning ones whose gains concentrate on frame-to-frame change while still retaining useful appearance.

To separate those factors, the authors run a matched study of 4 times 6 = 24 architecture-objective combinations. The study uses roughly 170M to 190M encoder scale and about 1.7M clips from OpenVid and Moments-in-Time v2, trained for 8 epochs.

On top of that study they propose TT-VidT. It pairs a DINOv3-initialized ViT-B/16 per-frame spatial path with a compact Temporal Transfer Layer. Training uses Diff Compression: the model learns to reconstruct target frames from a first-frame appearance anchor plus frame-specific motion tokens.

The sweep shows that TT3D combined with Diff Compression, and not either component alone, enters the strongest motion-sensitive regime. Decoder ablations favor a compact video-pretrained decoder.

In the final comparison, the authors report that TT-VidT leads Jester, Something-Something V2, ARID and Diving48 fine-tuning simultaneously. It improves over the strongest non-TT row by 54% to 121%. It does so with 48% fewer encoder FLOPs than DisMo and 55% fewer than VideoMAE or V-JEPA2. The abstract ends by noting that HMDB51, IARD and EPIC-Kitchens bound the claim.

Key facts

  • TT-VidT pairs a DINOv3-initialized ViT-B/16 per-frame spatial path with a compact Temporal Transfer Layer.
  • It is trained with Diff Compression, which reconstructs target frames from a first-frame appearance anchor and frame-specific motion tokens.
  • The design comes from a matched 4 times 6 = 24 architecture-objective study at roughly 170M to 190M encoder scale, on about 1.7M OpenVid and Moments-in-Time v2 clips for 8 epochs.
  • The authors report leading Jester, Something-Something V2, ARID and Diving48 fine-tuning, improving over the strongest non-TT row by 54% to 121%.
  • Encoder FLOPs are 48% lower than DisMo and 55% lower than VideoMAE or V-JEPA2; HMDB51, IARD and EPIC-Kitchens bound the claim.

Why it matters

Video pretraining comparisons are muddied when whole recipes are pitted against each other, so it is unclear what actually helps. This paper tries to fix that with a controlled 24-configuration sweep, and reports that the combination of TT3D and Diff Compression, rather than either piece alone, gives the strongest motion-sensitive representations. It also reports a lower compute bill: 48% fewer encoder FLOPs than DisMo and 55% fewer than VideoMAE or V-JEPA2.

Who it affects

Mainly researchers working on self-supervised video representation learning, especially those who care about motion rather than static appearance. Teams comparing video pretraining methods may also take note of the authors' point that architecture, objective, data, schedule, scale and decoder should be varied one at a time.

How to use it

No code, weights or release is mentioned. What can be taken from the abstract is the recipe: a DINOv3-initialized ViT-B/16 spatial path per frame, a compact Temporal Transfer Layer, Diff Compression as the objective, and a compact video-pretrained decoder, which the decoder ablations favor.

How solid is it

The evidence is a matched study of 24 architecture-objective combinations at a fixed scale, data budget and schedule, followed by a final comparison on seven benchmarks. The authors report leads on four of them: Jester, Something-Something V2, ARID and Diving48. The 54% to 121% range is relative improvement over the strongest non-TT row, and the abstract does not say which benchmark gets which value. Absolute scores are not given. All of this is the authors' own reporting in an abstract.

Risks and caveats

The authors themselves say HMDB51, IARD and EPIC-Kitchens bound the claim, and the abstract gives no detail on how. The study runs at roughly 170M to 190M encoder scale on about 1.7M clips for 8 epochs, so results at other scales are not covered by what is stated. The headline gains are relative percentages without absolute accuracies in the text.