WING moves human egocentric video skills to robot policies

WING moves human egocentric video skills to robot policies

The paper starts from a familiar bottleneck: learning general-purpose robot policies needs large-scale real-world interaction data, and collecting robot demonstrations remains expensive and difficult to scale. Egocentric videos, filmed from the human's own viewpoint, hold abundant human interaction experience with task-relevant semantics for manipulation. But the authors say direct transfer to robots is hard for two reasons. First, latent actions inferred from frame reconstruction can be dominated by nuisance variation such as ego-camera motion. Second, human and robot behaviors often exhibit different temporal dynamics.

Their answer is WING, short for World Action Learning via INteraction-Centric Spectral Latent Guidance. It is a framework for transferring interaction knowledge from egocentric videos to robot policies, and it works in two steps. WING first separates observer-induced motion from hand-object interaction and distills the interaction-centric component into latent actions. It then relies on the observation that cross-embodiment task semantics are concentrated in slowly varying temporal structures. WING identifies shared low-frequency components between the egocentric latent actions and robot behaviors in the spectral domain, and uses them to guide action generation.

On simulation benchmarks, WING achieves average success rates of 99.20% on LIBERO, 93.80% on RoboTwin 2.0 and 57.7% on RoboCasa-GR1. It also performs strongly across four real-world manipulation tasks under diverse generalization settings. From these results the authors conclude that interaction-centric spectral guidance provides an effective and scalable way to transfer physical interaction knowledge from human egocentric video to robot control. A project page is listed at https://mikuz12.github.io/wing/.

Key facts

  • WING (World Action Learning via INteraction-Centric Spectral Latent Guidance) transfers interaction knowledge from egocentric human videos to robot policies.
  • It separates observer-induced motion, such as ego-camera movement, from hand-object interaction, and distills the interaction part into latent actions.
  • It finds shared low-frequency components between egocentric latent actions and robot behaviors in the spectral domain and uses them to guide action generation.
  • Reported average success rates: 99.20% on LIBERO, 93.80% on RoboTwin 2.0 and 57.7% on RoboCasa-GR1, plus strong performance on four real-world manipulation tasks.

Why it matters

Robot demonstrations are expensive to collect and hard to scale, while egocentric human video is abundant. The obstacle is that raw human video does not map cleanly onto robot control: camera motion pollutes the inferred latent actions, and humans and robots move with different temporal dynamics. WING targets both problems at once, by isolating the hand-object interaction and by matching human and robot behavior through slowly varying, low-frequency structure. The authors argue that this is an effective and scalable route for moving physical interaction knowledge from human video into robot control.

Who it affects

The work speaks mainly to robot-learning researchers who want to train manipulation policies with less reliance on costly robot demonstrations, and to anyone building on egocentric human video as a data source. The benchmarks it reports on are LIBERO, RoboTwin 2.0 and RoboCasa-GR1.

How to use it

The paper gives a project page at https://mikuz12.github.io/wing/. The abstract does not say whether the code or models are released. Practitioners should treat WING for now as a method to read about and compare against, not a ready-made tool.

How solid is it

The reported numbers are average success rates of 99.20% on LIBERO, 93.80% on RoboTwin 2.0 and 57.7% on RoboCasa-GR1, which are the authors' own results. The abstract does not say what the benchmark success rates are averaged over (tasks, seeds or settings). No baseline methods or their scores are given, so no improvement over prior work can be stated. The real-world results are described only as strong across four manipulation tasks under diverse generalization settings; they are not quantified, and the four tasks and the robot hardware are not named.

Risks and caveats

The 57.7% on RoboCasa-GR1 is far below the other two benchmark figures, so results vary widely by benchmark. Without baselines, the numbers cannot show how much WING improves on earlier approaches. The real-world evidence is qualitative in the abstract. The central premise, that cross-embodiment task semantics sit in slowly varying temporal structures, is the authors' observation, and the abstract's own framing is that human and robot behaviors differ in temporal dynamics.

“interaction-centric spectral guidance provides an effective and scalable way to transfer physical interaction knowledge from human egocentric video to robot control”

— WING paper abstract