Latent-Foresight trains world model tokenizer and predictor jointly

Latent-Foresight trains world model tokenizer and predictor jointly

Predicting how a scene will evolve is a core capability for world models. The paper starts from recent work showing that operating in the feature space of Vision Foundation Models (VFMs) gives semantically rich representations that support a range of future scene understanding tasks.

The authors take issue with how existing approaches are built. They describe two-stage pipelines: VFM features are first compressed, either with fixed dimensionality reduction such as PCA or with independently trained autoencoders, and then a separate predictor is trained on top of the resulting frozen latent space. The authors argue that this split between representation learning and temporal prediction, like approaches that apply predictors directly to raw VFM features, provides no guarantee that the latent space is structured for predictable dynamics.

Their answer is Latent-Foresight, an end-to-end framework that jointly learns a latent tokenizer and a flow-based generative dynamics model. Because both parts are trained together, the representation is explicitly shaped to support temporal predictability rather than being fixed before the predictor ever sees it.

Joint training is not trivially stable, so the authors introduce several design choices that prevent latent collapse and align reconstruction with generative objectives. The abstract does not describe what those choices are.

The experiments, according to the authors, show that the approach learns more temporally coherent latent representations and consistently outperforms two-stage baselines across multiple future scene understanding tasks and prediction horizons. It also removes the separate training stages, including during high-resolution adaptation. The authors provide implementation code and model weights on GitHub at github.com/Sta8is/Latent-Foresight.

Key facts

  • Latent-Foresight jointly learns a latent tokenizer and a flow-based generative dynamics model on top of Vision Foundation Model features.
  • It targets two-stage pipelines, where VFM features are compressed with fixed reduction such as PCA or separate autoencoders and a predictor is trained on the frozen latent space.
  • Design choices that prevent latent collapse and align reconstruction with generative objectives are introduced to keep joint training stable.
  • The authors report consistent gains over two-stage baselines across multiple future scene understanding tasks and prediction horizons, and no separate training stages, including during high-resolution adaptation.
  • Code and model weights are provided at github.com/Sta8is/Latent-Foresight.

Why it matters

World models need to forecast how a scene changes, and many recent systems do this in the feature space of Vision Foundation Models. The paper's point is that compressing those features first and predicting afterwards gives no guarantee the latent space is easy to predict. Learning the representation and the dynamics together is a direct attempt to fix that, and it also collapses several training stages into one.

Who it affects

Researchers building latent world models and future scene understanding systems on top of Vision Foundation Model features are the obvious audience, especially those who currently use PCA or a separately trained autoencoder before a predictor.

How to use it

The authors provide implementation code and model weights at github.com/Sta8is/Latent-Foresight, so the method can be inspected and tried directly.

How solid is it

The claims come from the authors' own abstract: they say extensive experiments show more temporally coherent latent representations and consistent gains over two-stage baselines across multiple tasks and prediction horizons. No quantitative results are given in the abstract. The released code and weights make independent checking possible.

Risks and caveats

The abstract gives no metrics, margins, named baselines, datasets or tasks, so the size of the improvement is unknown from this text. The design choices that prevent latent collapse are not described, and joint training of this kind needs them to stay stable. Model size, training compute and the license of the code and weights are not stated.