LOCI pairs a key-value cache with linear memory for streaming video world models

LOCI pairs a key-value cache with linear memory for streaming video world models

A paper on Hugging Face introduces LOCI, a hybrid spatial-memory architecture for streaming video world models. It targets one specific problem: when a camera revisits a region it has already seen, the model should reproduce what was there before. That takes two things, remembering past observations and retrieving the right one for the current viewpoint.

The paper frames the trade-off between the two usual approaches. Key-value caches preserve visual detail, but they grow with video length. Recurrent memory is compact, but it compresses history into a fixed-size state, so individual past observations can no longer be accessed directly. LOCI keeps both representations.

The split works at the level of transformer blocks. In half of the blocks, main attention keeps a key-value cache of past observations. In the other half, attention is restricted to the current chunk and complemented by a recurrent linear-attention memory. The reads and writes of that recurrent memory are conditioned on projective camera geometry, so viewpoint enters both the addressing of memory and the content stored in it. Recurrent readouts then flow into the subsequent cache-backed blocks and give their queries accumulated scene context.

The authors report results on the public MIND memory benchmark and on held-out recorded trajectories. There, LOCI reproduces revisited content more faithfully than representative world models and than a full-softmax model trained with the same recipe. With full history, it lowers peak memory at equal length by about 30% relative to full softmax. With a bounded bank of retained observations, it streams long videos at constant memory and stays more faithful than full softmax under the same budget.

Key facts

  • LOCI is a hybrid spatial-memory architecture for streaming video world models that keeps both a key-value cache and a recurrent linear-attention memory.
  • In half of the transformer blocks, main attention keeps a key-value cache of past observations; in the other half, attention is restricted to the current chunk and complemented by recurrent memory.
  • The recurrent memory's reads and writes are conditioned on projective camera geometry, so viewpoint shapes both addressing and stored content.
  • On the public MIND memory benchmark and held-out recorded trajectories, LOCI reproduces revisited content more faithfully than representative world models and a same-recipe full-softmax model.
  • With full history, peak memory at equal length is about 30% lower than full softmax; with a bounded bank of retained observations, it streams long videos at constant memory.

Why it matters

Video world models have to stay consistent when a camera returns to a place it has already shown. The two standard memory choices each fail in one direction: key-value caches keep visual detail but grow with video length, while recurrent memory stays compact but squeezes history into a fixed-size state. LOCI tries to get the benefits of both by splitting the work across transformer blocks and tying memory access to camera geometry.

Who it affects

The work is aimed at researchers building streaming video world models, especially those who care about long videos and scene consistency when a camera revisits earlier viewpoints. The summary does not name authors or institutions.

How to use it

There is nothing to run yet from what the summary describes. It mentions no code or model release. The practical takeaway is architectural: keep a key-value cache in some blocks, use a camera-conditioned recurrent linear-attention memory in the others, and feed the recurrent readouts into the cache-backed blocks as extra context for their queries.

How solid is it

The claims come from the authors' own summary of their results. They cite the public MIND memory benchmark and held-out recorded trajectories, and a comparison against a full-softmax model built with the same recipe. The comparison with other world models is stated only qualitatively (more faithful), and the other world models are not named. No numeric faithfulness scores are given. The one figure is the roughly 30% peak-memory reduction versus full softmax, which applies to the full-history setting.

Risks and caveats

Several details are missing: model size, training data, the size of the bounded bank, and the video lengths tested. The 30% memory saving is stated for the full-history setting only; for the bounded-bank setting the source says only that memory is constant. Without numeric scores, it is hard to judge how large the faithfulness gains are.