DreamTrue robot world model takes first place in AgiBot World Challenge 2026

DreamTrue robot world model takes first place in AgiBot World Challenge 2026

DreamTrue is a multi-view, cross-embodiment robot world model built for action-faithful and physically plausible video prediction: given a robot's actions, it generates video of what happens next, and the video should both follow those actions and obey physics.

The authors say that training such a model on existing robot datasets runs into two obstacles. Imprecise calibration can impair action following, while limited coverage of unsuccessful interactions can bias predictions toward successful outcomes. The paper answers each problem, then adds a third piece for feedback.

First, to improve action following across embodiments, the authors render action trajectories into image-space conditions. They introduce offline geometric calibration to align these conditions with the target videos.

Second, to broaden interaction coverage, they introduce counterfactual post-training. It modifies recorded action trajectories and generates future videos under a wider range of actions and contact configurations, so the model sees more than the successful outcomes in the recorded data.

Third, to give feedback on predictions when there are no paired ground-truth futures, the authors construct a human-annotated video dataset covering robot, object and interaction defects. They use it to train an embodied video reward model. Its scores guide reinforcement-learning post-training toward more physically plausible interaction outcomes.

On AgiBot, the authors report that DreamTrue attains state-of-the-art action following, while reducing the human-assessed interaction defect rate from 48.12% to 6.25%. They also state that the model ranks first in the world model track of the AgiBot World Challenge 2026. The project page is at https://brave-eai.github.io/DreamTrue.

Key facts

  • DreamTrue is a multi-view, cross-embodiment robot world model for action-faithful, physically plausible video prediction.
  • Offline geometric calibration aligns image-space renderings of action trajectories with the target videos to improve action following across embodiments.
  • Counterfactual post-training modifies recorded action trajectories and generates future videos under a wider range of actions and contact configurations, covering more unsuccessful interactions.
  • A reward model trained on a human-annotated dataset of robot, object and interaction defects guides reinforcement-learning post-training.
  • On AgiBot, the authors report state-of-the-art action following and a human-assessed interaction defect rate down from 48.12% to 6.25%, plus first place in the world model track of the AgiBot World Challenge 2026.

Why it matters

A world model that predicts what a robot's actions will cause is only useful if the predicted video actually follows those actions and respects physics. The authors name two weaknesses of models trained on existing robot data: imprecise calibration hurts action following, and a shortage of failed interactions pushes predictions toward success. DreamTrue targets both, using geometric calibration for the first and counterfactual post-training for the second. A video reward model then supplies feedback where no ground-truth future exists.

Who it affects

The work is aimed at researchers building robot world models, especially models that must work across different robot embodiments. It also concerns participants in the AgiBot World Challenge 2026, where DreamTrue reports first place in the world model track.

How to use it

The abstract points to a project page at https://brave-eai.github.io/DreamTrue. No code or weights release is stated; only a project page is mentioned. For practitioners, the transferable ideas are the three stages described: calibrate action conditions offline, post-train on counterfactually modified trajectories, and post-train with reinforcement learning guided by a defect-aware video reward model.

How solid is it

The claims come from the paper's abstract and are the authors' own. The headline result is a human-assessed interaction defect rate falling from 48.12% to 6.25% on AgiBot, alongside state-of-the-art action following and a first-place ranking in the world model track of the AgiBot World Challenge 2026. No numeric action-following metric is given, only the claim of state-of-the-art. It is not stated which baseline the 48.12% defect rate belongs to (for example the model before RL post-training or a prior model), nor how many samples were assessed.

Risks and caveats

Several details that would let a reader judge the result are not given: no model size, architecture details, training data volume or compute; the number of annotated videos or of human raters is not given. No limitations or real-robot deployment results are mentioned. The defect-rate figure depends on human assessment, and the baseline behind the 48.12% starting value is not identified, so the size of the improvement cannot be placed against a specific prior model.

“imprecise calibration can impair action following, while limited coverage of unsuccessful interactions can bias predictions toward successful outcomes”

— DreamTrue paper abstract