Long-WAM lifts robot success with 19.2 seconds of video context

Long-WAM lifts robot success with 19.2 seconds of video context

Long-WAM is a model-system framework for scaling the context of causal world-action models under real-time robot control constraints. The problem it targets: real-time control needs enough visual history to infer motion and task progress, but processing that history can delay action.

The authors' central finding is that access to history is not the same as using it. Longer histories pay off far more when the video foundation is pretrained autoregressively (AR). Their training recipe first learns causal prediction from robot and egocentric videos without action labels, then preserves this history-to-future structure during world-action adaptation.

The headline benchmark result is on RoboCasa GR-1. Increasing context from 0.0 to 19.2 seconds raises success from 63.3% to 78.7%, a gain of 15.4 percentage points. A bidirectionally pretrained initialization shows no net gain from the same extra context. Robot-domain AR pretraining raises peak success further on GR-1 and LIBERO-Long. The authors also report that Long-WAM achieves the best results among compared methods on LIBERO-Long, RoboTwin 2.0 and DOMINO.

On the systems side, streaming observation encoding, asynchronous execution and hardware-specific acceleration let the model run on an RTX 5090, a DGX Spark and a Jetson AGX Thor without dropping future prediction. On the RTX 5090, each action chunk, including future-video latent prediction, takes 107.4 ms.

On real robots, the authors say deployment on Unitree G1 and YAM supports dynamic and long-horizon manipulation. The standout number is 95% success on dynamic cup stacking, a task where Pi0.5 and Fast-WAM each succeed in none of 20 trials. The authors add that, as a memory-informed executor, Long-WAM also complements higher-level planning in composite tasks.

Key facts

  • On RoboCasa GR-1, raising context from 0.0 to 19.2 seconds lifts Long-WAM's success from 63.3% to 78.7%; a bidirectionally pretrained initialization shows no net gain.
  • The central finding: longer visual history pays off far more when the video foundation is pretrained autoregressively.
  • On an RTX 5090, each action chunk, including future-video latent prediction, takes 107.4 ms; the system also targets DGX Spark and Jetson AGX Thor.
  • On real robots (Unitree G1 and YAM), Long-WAM reaches 95% success on dynamic cup stacking, while Pi0.5 and Fast-WAM succeed in none of 20 trials.
  • The authors report the best results among compared methods on LIBERO-Long, RoboTwin 2.0 and DOMINO.

Why it matters

Robots that act in real time need to remember what just happened, but feeding in long video history can slow the controller down. Long-WAM tackles both sides of that tension. Its reported lesson is that simply giving a model more history does not guarantee it will use it: the same extra context helped an autoregressively pretrained model on RoboCasa GR-1 (63.3% to 78.7%) and gave no net gain to a bidirectionally pretrained one. That points to the pretraining objective, not only the context window, as what decides whether memory helps.

Who it affects

Mainly robotics researchers building world-action models and video-based robot policies, who get a pretraining recipe and a set of comparisons on RoboCasa GR-1, LIBERO-Long, RoboTwin 2.0 and DOMINO. The deployment work on RTX 5090, DGX Spark and Jetson AGX Thor, and on Unitree G1 and YAM robots, is relevant to teams who want to run such models on onboard or desktop-class hardware.

How to use it

The abstract describes a two-stage recipe: first learn causal prediction from robot and egocentric videos without action labels, then adapt to world-action modelling while preserving the history-to-future structure. For deployment it names streaming observation encoding, asynchronous execution and hardware-specific acceleration. The abstract does not say whether code, weights or data are released.

How solid is it

This is a research paper whose abstract is the basis for this account. All figures are the authors' own reports, and the abstract names no authors or institutions. The 63.3% to 78.7% result is stated for RoboCasa GR-1 only; the abstract gives no numeric success rates for LIBERO-Long, RoboTwin 2.0 or DOMINO. The abstract gives no success rate for the bidirectionally pretrained initialization, only that it shows no net gain. The number of trials behind the 95% cup-stacking result is not stated explicitly for Long-WAM; 20 trials is stated for Pi0.5 and Fast-WAM.

Risks and caveats

The headline gain comes from one benchmark, RoboCasa GR-1, and the best-among-compared claims on other benchmarks come without numbers in the abstract. The real-robot comparison is reported for a single task, dynamic cup stacking. No latency figures are given for DGX Spark or Jetson AGX Thor, and the model size (parameter count) is not stated.

“access to history is not the same as using it: longer histories pay off far more when the video foundation is pretrained autoregressively (AR)”

— Long-WAM paper abstract