World Embedding Benchmark tests how video embeddings encode physics

World Embedding Benchmark tests how video embeddings encode physics

Physical fidelity has drawn growing attention in world models and video generation, but how video representations encode physical information is less well understood. A new paper introduces the World Embedding Benchmark to probe exactly that.

The benchmark has 8,000 controlled simulation cases drawn from 80 families. They span four areas: fluid mechanics, solid mechanics, dynamics, and optics and electromagnetism. Each case pairs a rendered video with physical annotations derived from the simulation. Three complementary tasks run on top of this data: text-video retrieval, physical-property regression, and multiple-choice video-description pair classification. The authors use the tasks to separate two things: cross-modal physical alignment, and the recoverability of quantitative physical information.

The results come in three parts. First, the evaluated pre-trained omnimodal embedding models show weak retrieval and near-chance within-family pair classification. Yet lightweight probes can still recover useful physical information from the frozen video embeddings, so the information is present even when alignment is poor.

Second, the authors tried continual contrastive training with physics-specific video-text pairs. It improves retrieval and pair classification but degrades physical-property regression. They read this as a trade-off between alignment and how well quantitative information can be recovered.

Third, they used the embeddings to retrieve reference videos for retrieval-augmented generation with MiniMax-H3. In their experiments, the retrieved references improve the physical fidelity of generated videos, and stronger retrieval models give larger gains.

The authors conclude that physical alignment and property recoverability should be evaluated jointly, and that physical representations are useful for improving video generation.

Key facts

  • The World Embedding Benchmark has 8,000 controlled simulation cases from 80 families, covering fluid mechanics, solid mechanics, dynamics, and optics and electromagnetism.
  • It supports three tasks: text-video retrieval, physical-property regression, and multiple-choice video-description pair classification.
  • Evaluated pre-trained omnimodal embedding models show weak retrieval and near-chance within-family pair classification, while lightweight probes recover useful physical information from frozen embeddings.
  • Continual contrastive training on physics-specific video-text pairs improves retrieval and pair classification but degrades physical-property regression.
  • In the authors' experiments, retrieved reference videos improve the physical fidelity of MiniMax-H3 generations, with stronger retrieval models giving larger gains.

Why it matters

Video generation and world models are judged more and more on whether they respect physics, yet the representations underneath are poorly understood. This benchmark gives a way to ask the question directly. Its main finding is a split between two properties: an embedding can line up well with text and still lose quantitative physical detail, and the reverse. Training that improved retrieval and pair classification made property regression worse, so a single alignment score can hide a loss. The authors argue the two should be evaluated jointly.

Who it affects

Researchers building or evaluating video embeddings, world models and physics-aware video generators are the direct audience. Teams that use retrieval to condition video generation also have a stake, since the paper reports that better retrievers gave larger gains in physical fidelity in the authors' experiments.

How to use it

The benchmark's three tasks can serve as a template for testing a video embedding model: retrieval against text, regression of physical properties from frozen embeddings with lightweight probes, and multiple-choice pair classification. Retrieving reference videos to feed a generator, as the authors did with MiniMax-H3, is the downstream use they demonstrate. No release, dataset link, code availability or license is stated in the abstract.

How solid is it

The design is clear: simulation-derived annotations give controlled ground truth across 80 families and four physics areas. The findings are attributed to the authors' own experiments. No numeric scores, accuracy figures or sizes of gains are given for retrieval, classification, regression or video generation. Which embedding models were evaluated is not named; only 'pre-trained omnimodal embedding models' is stated. The abstract names no authors or institutions.

Risks and caveats

'Near-chance' is not quantified; no chance level is given. The gains from retrieval are stated only 'in our experiments'; no claim of generality is made. No information is given on what MiniMax-H3 is, who makes it, or how fidelity of generated videos was measured. The trade-off between alignment and regression is reported for one training approach, continual contrastive training with physics-specific video-text pairs.

“these findings highlight the need to evaluate physical alignment and property recoverability jointly”

— The authors, World Embedding Benchmark abstract