H2R-Bench finds video world models struggle with human-to-robot transfer

H2R-Bench finds video world models struggle with human-to-robot transfer

Robot learning needs large amounts of manipulation data, but collecting real robot demonstrations is expensive and hard to scale. Egocentric video of humans performing manipulation tasks is abundant by comparison, but transferring that behavior to a robot is difficult because human hands and robotic end-effectors move differently. Video world models have recently offered a promising path around this: synthesizing robot-centric manipulation video from human observations. Whether they can actually make that cross-embodiment jump had remained largely unexplored.

To test this, the authors introduce H2R-Bench, a benchmark for cross-embodiment human-to-robot manipulation video generation: a model is given an egocentric human demonstration video and a target embodiment, and must generate the corresponding robot manipulation video. Each benchmark instance pairs a human demonstration video with target embodiment constraints and source-grounded annotations covering task goals, action events, functional contacts, and object responses. Generated videos are scored on five dimensions: goal-state completion, action-event completion, functional contact transfer, embodiment correctness, and general video quality.

The authors ran eleven state-of-the-art video generation models through H2R-Bench, spanning six manipulation families and two robot embodiments. Current video world models remain limited in this kind of transfer: even the leading models often fail in embodiment consistency, functional interaction, and task execution. The abstract reports this pattern qualitatively; it does not give numeric scores, pass rates, or rankings for how individual models performed on the five dimensions.

The authors frame H2R-Bench as a systematic diagnostic framework for evaluating whether video world models can bridge the human-to-robot embodiment gap and convert human manipulation observations into training resources for robots.

Key facts

  • H2R-Bench evaluates cross-embodiment human-to-robot manipulation video generation: a model is given an egocentric human demonstration video and a target embodiment, and must generate the matching robot manipulation video.
  • Each benchmark instance pairs a human demonstration video with target embodiment constraints and source-grounded annotations of task goals, action events, functional contacts, and object responses.
  • Generated videos are scored on five dimensions: goal-state completion, action-event completion, functional contact transfer, embodiment correctness, and general video quality.
  • Eleven state-of-the-art video generation models were benchmarked across six manipulation families and two robot embodiments.
  • Even the leading models often fail in embodiment consistency, functional interaction, and task execution, showing current video world models remain limited in human-to-robot transfer.

Why it matters

Robot learning is bottlenecked by data: real robot demonstrations are expensive and hard to collect at scale, while human manipulation video is abundant. Video world models offer a possible way to turn that human footage into synthetic robot training video, but whether these models can actually make the jump from a human hand to a robot end-effector had remained largely unexplored. H2R-Bench gives that question a systematic evaluation across eleven models and five distinct dimensions of video quality and correctness.

Who it affects

Researchers building or evaluating video world models for robotics, and robot-learning teams weighing synthetic, human-derived video as a substitute for costly real robot demonstrations. The five-dimension scoring, covering goal-state completion, action-event completion, functional contact transfer, embodiment correctness, and general video quality, gives both groups a shared way to measure this specific transfer capability rather than general video generation quality.

How to use it

Each H2R-Bench instance pairs a human demonstration video with a target embodiment and source-grounded annotations of task goals, action events, functional contacts, and object responses, so a generated video can be checked against a specific target rather than judged on how realistic it looks. Teams developing video world models for robotics can use it to locate exactly where cross-embodiment transfer breaks down, in embodiment consistency, functional interaction, or task execution, before treating generated video as usable robot training data.

How solid is it

The evaluation is broad: eleven state-of-the-art video generation models tested across six manipulation families and two robot embodiments, scored on five separate dimensions rather than a single aggregate quality score, with each instance's targets grounded in annotations of the human demonstration rather than assigned after the fact. What the abstract does not provide is a breakdown: it reports the failure pattern qualitatively and does not give numeric scores, pass rates, or rankings for how individual models performed on the five dimensions.

Risks and caveats

The paper's own finding is the caveat: current video world models remain limited in human-to-robot manipulation transfer, and even the leading models among the eleven tested often fail in embodiment consistency, functional interaction, and task execution. That limitation applies broadly, across six manipulation families and two robot embodiments, not to an isolated case.