Honeycomb video world model keeps scene memory at constant size

Honeycomb video world model keeps scene memory at constant size

Video world models need a persistent memory of the scene to stay consistent over long-horizon video generation. The authors of a new paper argue that existing spatial memories accumulate RGB observations or latent features, so storage requirements keep rising as generation proceeds.

Their answer is Honeycomb, a video world model built on HexMemory. HexMemory is a low-rank representation that stores scene features in a fixed-size memory made up of a total of six spatial and spatiotemporal planes.

The system works in three parts. A feed-forward writer maps each generated chunk of video into new plane features. When the spatial coverage or temporal range expands, the previous planes are warped while preserving their dimensions, then fused with the new features through confidence-weighted pooling and a learned residual correction. A reader then retrieves latents from HexMemory to condition the next stretch of video generation.

Because the writer processes only observations from the new chunk, it avoids per-scene optimization and repeated processing of the full history.

The authors report experiments on WorldScore and RealEstate10K. According to them, these show strong video generation quality and robust revisit consistency, while HexMemory feature storage stays constant throughout generation. Code and additional visualizations are available on the project page at https://jackswl.github.io/honeycomb/.

Key facts

  • Honeycomb is a video world model built on HexMemory, a low-rank representation proposed by the authors.
  • HexMemory stores scene features in a fixed-size memory of a total of six spatial and spatiotemporal planes.
  • A feed-forward writer processes only the new chunk; old planes are warped, then fused using confidence-weighted pooling and a learned residual correction.
  • A reader retrieves latents from HexMemory to condition later generation.
  • Experiments on WorldScore and RealEstate10K are reported to show strong quality and robust revisit consistency with constant feature storage; code is on the project page.

Why it matters

Long-horizon video generation needs the model to remember what a scene looks like so it can return to a place and keep it consistent. The authors say existing spatial memories grow because they accumulate RGB observations or latent features. Honeycomb's pitch is a memory whose size stays fixed however long generation runs.

Who it affects

Mainly researchers working on video world models and long-horizon video generation, who deal with scene consistency and memory cost. The paper evaluates on WorldScore and RealEstate10K.

How to use it

The authors say code and additional visualizations are available on their project page at https://jackswl.github.io/honeycomb/. The text does not describe setup steps.

How solid is it

The account rests on the paper's abstract, so these are the authors' own claims. They say experiments on WorldScore and RealEstate10K show strong generation quality and robust revisit consistency with constant HexMemory feature storage. The abstract gives no scores, baselines or comparisons to named methods, so the size of any improvement cannot be judged from it.

Risks and caveats

The abstract gives no numbers for quality or consistency, no plane dimensions or memory size, and no generation length, so the practical limits of a six-plane memory are unknown from this text. Claims of strong quality are the authors' own and are not independently checked here.