WorldGuide, a closed-loop video world model, reports higher task success than MiniMax-H3

A paper on Hugging Face introduces WorldGuide, a video world model built for long procedural tasks. The authors start from a familiar limit: video generators and video-based world models can produce plausible visual trajectories, but a long task needs generation that adapts to what has actually been produced. The model must pick the next action from its generated state, carry it out, and recognise when the task is done. Open-loop generation cannot adapt to execution outcomes, the authors say, while existing closed-loop systems often rely on pretrained executors or indirect verification. That leaves a gap between deciding an action and actually realising it.
WorldGuide frames procedural video generation as closed-loop task execution in visual world space. It is given only an initial image and a task goal. It predicts an atomic action, generates the matching video clip, and then uses that generated result to select the next action or to terminate. The system has two parts, a Planner and an Executor, trained on the same step-level procedural demonstrations. The Planner learns to predict the next atomic action, or task completion, from visual progress. The Executor is trained directly to realise the actions the Planner predicts. A hierarchical visual memory keeps state across long-horizon execution while holding the cost of history tokens bounded.
Because step-level action-video supervision for joint planner-executor training is lacking, the authors also built WorldGuide Bench: approximately 59K step-annotated videos across 245 tasks and 27 procedural categories.
On results, WorldGuide achieves 33.33% Task Success on WorldGuide-Bench, against 29.90% for the strong recent video model MiniMax-H3, even though MiniMax-H3 receives reference action plans. That is a gap of 3.43 percentage points. On VideoCraft-Bench, WorldGuide achieves 47.69% against 32.73% for MiniMax-H3 under goal-only conditioning, a gap of 14.96 percentage points. The authors conclude that these results demonstrate the importance of coupling planning with learned execution for goal-directed procedural video generation.
Key facts
- WorldGuide takes an initial image and a task goal, predicts an atomic action, generates its video clip, and uses the result to pick the next action or stop.
- A Planner and an Executor are trained on the same step-level procedural demonstrations, with a hierarchical visual memory that bounds the history token cost.
- The authors introduce WorldGuide Bench: approximately 59K step-annotated videos across 245 tasks and 27 procedural categories.
- On WorldGuide-Bench, WorldGuide reaches 33.33% Task Success versus 29.90% for MiniMax-H3, which was given reference action plans.
- On VideoCraft-Bench, WorldGuide reaches 47.69% versus 32.73% for MiniMax-H3 under goal-only conditioning.
Why it matters
Most video generation is open loop: the model produces a clip without checking what it actually produced. WorldGuide tries to close that loop for procedural tasks by making each generated clip the input for deciding what happens next, including when to stop. The authors present this as a way to narrow the gap between deciding an action and successfully realising it. The headline claim is that a model given only a goal beat a strong recent video model that was handed reference action plans on one benchmark.
Who it affects
Mainly researchers working on video generation, video-based world models and goal-directed or procedural video. The work is about generated video, and the source gives no information on real-world or robotic deployment. The new WorldGuide Bench, with about 59K step-annotated videos, targets a data gap the authors say exists for joint planner-executor training.
How to use it
There is nothing to use directly yet: no code, model weights or dataset release is mentioned. What a reader can take away is the design. A Planner that predicts the next atomic action or completion from visual progress, an Executor trained on the same step-level demonstrations to realise those actions, and a hierarchical visual memory to keep long-horizon state affordable. The benchmark composition (245 tasks, 27 procedural categories) is described, but access to it is not.
How solid is it
These are the authors' own results on a benchmark they built and on VideoCraft-Bench, compared against a single named baseline, MiniMax-H3. The abstract names no authors or institutions. The source does not define how Task Success is computed or judged, and no details on VideoCraft-Bench (size, origin) are given. The WorldGuide-Bench margin is 3.43 points; the VideoCraft-Bench margin is 14.96 points.
Risks and caveats
The two comparisons use different conditions for the baseline: MiniMax-H3 received reference action plans on WorldGuide-Bench and goal-only conditioning on VideoCraft-Bench, so the two gaps are not like for like. The smaller margin sits on the authors' own benchmark. Beyond the two Task Success comparisons, no ablations or other baselines are reported, and no model size, training compute or inference cost is given. The conclusion about coupling planning with learned execution is the authors' reading of these two numbers.
“These results demonstrate the importance of coupling planning with learned execution for goal-directed procedural video generation.”
— WorldGuide paper abstract