OmniVBench benchmarks omni reference-to-video generation with 12,172 checklist items

OmniVBench benchmarks omni reference-to-video generation with 12,172 checklist items

Reference-to-video (R2V) generation, where a model produces a video guided by one or more reference inputs, is moving toward a more general and versatile form of control that the paper's authors call omni R2V. They argue existing benchmarks have not kept pace: current test suites cover only a limited range of reference types and combinations, and their evaluation protocols mostly judge whether a generated video is holistically consistent with its references, without checking whether individual reference factors, such as content, motion or style, are properly preserved, kept separate from each other, and correctly routed to the right part of the output. Building suitable training data for omni R2V is also expensive, which the authors say has left the field short of adequate training resources.

To close both gaps, the authors introduce two things: OmniVBench, an evaluation benchmark, and the Omni-R2V Dataset, a training dataset. OmniVBench broadens R2V evaluation across a wider range of reference types, more fine-grained control tasks, and richer combinations of references, spanning 7 task families and 18 fine-grained tasks that cover content, motion, style, structure, narrative, and multi-reference settings. Its core method is what the authors call factor-grounded evaluation: 12,172 case-specific checklist items that check, case by case, whether the reference factors a task calls for are faithfully preserved, correctly disentangled and bound to their intended targets, and properly realized according to the instruction, rather than just checking overall similarity to the references.

The Omni-R2V Dataset supplies industrial-grade training resources for a range of R2V tasks. Drawn mainly from a large-scale corpus of professional video footage, it comprises 340,000 processed training samples covering diverse reference types and multi-reference compositions. The authors also built task-specific pipelines for constructing reference-target pairs, which they present as a practical, scalable recipe for building omni R2V training data more broadly.

Using OmniVBench, the authors ran an extensive evaluation of advanced open- and closed-source R2V models. The evaluation found clear performance gaps across task families and across the benchmark's evaluation dimensions, which the authors say highlights remaining limitations of current R2V models. The abstract does not name the specific models tested or give numeric scores for the gaps found.

Key facts

  • OmniVBench evaluates omni reference-to-video (R2V) generation across 7 task families and 18 fine-grained tasks covering content, motion, style, structure, narrative and multi-reference settings.
  • Its factor-grounded evaluation uses 12,172 case-specific checklist items to check whether reference factors are preserved, disentangled and correctly bound to their targets, not just holistically consistent.
  • The accompanying Omni-R2V Dataset provides 340,000 processed training samples, drawn mainly from professional video footage, for training R2V models.
  • The authors built task-specific pipelines for constructing reference-target pairs as a scalable recipe for building omni R2V training data.
  • Evaluating advanced open- and closed-source R2V models on OmniVBench revealed clear performance gaps across task families and evaluation dimensions, without specific model names or scores given in the abstract.

Why it matters

R2V generation is moving toward controlling video output with an increasingly general and versatile set of references rather than a single fixed type. The authors argue that as this omni R2V paradigm emerges, the tools used to measure and train it have not kept up: benchmarks test a narrow slice of reference types and only check whether output looks broadly consistent with references, not whether each individual reference factor is actually preserved and kept distinct. OmniVBench and the Omni-R2V Dataset are offered as a matched pair of instruments, one for evaluation and one for training, meant to close that gap at once.

Who it affects

Researchers and engineering teams building or evaluating reference-to-video generation models are the direct audience: OmniVBench gives them a way to test omni R2V capability that existing benchmarks did not cover, and the Omni-R2V Dataset gives them training data for tasks where suitable resources were previously scarce. The paper frames the dataset as being handed to the broader research community rather than kept internal.

How to use it

OmniVBench is meant for evaluating R2V models across 7 task families and 18 fine-grained tasks, applying the factor-grounded checklist method (12,172 case-specific items) to score whether reference factors are preserved, disentangled and correctly bound. The Omni-R2V Dataset, with its 340K processed training samples and task-specific reference-target pair construction pipelines, is meant as a training resource for a diverse set of R2V tasks. The abstract does not state a release date, venue, or licensing terms for either the benchmark or the dataset.

How solid is it

The authors report an extensive evaluation of advanced open- and closed-source R2V models using OmniVBench, finding clear performance gaps across task families and evaluation dimensions. This is presented as evidence that the benchmark surfaces real, measurable weaknesses in current models rather than being a purely theoretical construct. The abstract does not name which models were evaluated or give the numeric scores behind the reported gaps, so the size and specifics of those gaps cannot be checked from the abstract alone.

Risks and caveats

The abstract does not disclose the provenance or licensing of the large-scale corpus of professional video footage the training dataset is drawn from, which matters for anyone assessing reuse rights. It also omits the full author list, institutional affiliations, a release date or venue, and the identities of the specific R2V models evaluated, along with their numeric scores. These are gaps in what the abstract states, not claims that the paper omits them entirely.