Ego2Act benchmark finds video models skip steps in multi-step tasks

Video generation models are increasingly being explored as world simulators for embodied planning and learning. For that to work, the paper says, a model must do more than produce attractive frames: it has to predict how an environment changes as goal-directed actions are carried out. The authors argue that existing benchmarks focus mainly on single short actions or step-by-step instructions, which leaves multi-step physical reasoning underexplored. That gap is sharpest in egocentric video generation, where a model has to plan and simulate the proper execution of several real-world manipulations to reach a high-level goal.\n\nTo address it, the authors introduce Ego2Act, a goal-directed benchmark with 2,640 videos drawn from 110 real-world tasks. The tasks span day-to-day settings and vary in object clutter and multi-step complexity. Each test gives a model an initial scene image and a high-level goal, and asks whether it can produce a realistic egocentric video of a hand manipulating objects to carry out the task.\n\nThe authors also introduce Ego2ActJudge, a reference-free evaluation pipeline meant to make this kind of evaluation scalable. They report that it aligns better with human consensus on task completion and physics plausibility than relevant baselines.\n\nTwo findings come out of the evaluation. First, the models' generated simulations often skip or partially execute steps, so later steps are left without the states they depend on, and the goal goes unfulfilled. Second, the models consistently fail at fine-grained physical dynamics, particularly during complex object manipulation and persistent world modeling. The authors close by saying they hope Ego2Act becomes a rigorous testbed for moving video models toward physically plausible, goal-directed simulation.
Key facts
- Ego2Act is a goal-directed benchmark with 2,640 videos from 110 real-world tasks, covering day-to-day settings with varying object clutter and multi-step complexity.
- Each test gives a model an initial scene image and a high-level goal, then checks whether it can generate a realistic egocentric video of a hand manipulating objects to complete the task.
- Ego2ActJudge is a reference-free evaluation pipeline that, per the authors, aligns better with human consensus on task completion and physics plausibility than relevant baselines.
- Finding one: generated simulations often skip or partially execute steps, leaving later steps missing dependent states and the goal unfulfilled.
- Finding two: models consistently fail at fine-grained physical dynamics, especially in complex object manipulation and persistent world modeling.
Why it matters
Video models are being explored as world simulators for embodied planning and learning, and that use needs more than good-looking frames. The authors say existing benchmarks mostly test single short actions or step-by-step instructions, so whether a model can carry out a whole multi-step goal has been underexplored. Ego2Act targets exactly that: given a scene and a high-level goal, can the model show the full task being done from a first-person view?
Who it affects
Researchers building or evaluating video generation models, especially those who want to use them as world simulators for embodied planning and learning. The findings are also a signal for anyone relying on generated video to show how an environment evolves under a sequence of actions.
How to use it
Ego2Act is a benchmark for testing a video model, not a product. The setup is simple to describe: supply an initial scene image and a high-level goal, generate an egocentric video, then score task completion and physics plausibility, with Ego2ActJudge offered as a scalable, reference-free way to do the scoring. No release date, code or dataset link is mentioned in the source.
How solid is it
The material is a paper abstract, so the claims are the authors' own. The benchmark size is stated concretely (2,640 videos, 110 real-world tasks). The claim that Ego2ActJudge beats baselines on alignment with human consensus is stated without figures: the size of the improvement is not stated, and the baselines are not named. No quantitative results are given, and no names of the evaluated video generation models are given.
Risks and caveats
The headline findings, skipped or partial steps and weak fine-grained physics, are reported in general terms. The abstract does not say which models fail more or less, and no scores or success rates are given, so the size of the problem cannot be judged from this text alone. The judge's reliability rests on the authors' own comparison with baselines, which should be checked against the full paper.
“We hope Ego2Act provides a rigorous testbed for advancing video models toward physically plausible, goal-directed simulation.”
— Ego2Act authors