4Director controls video world models with rigid 3D object meshes

4Director controls video world models with rigid 3D object meshes

A new paper introduces 4Director, a video world model aimed at precise control over camera and object motion, which the authors call essential for professional video production. They argue that existing methods control objects only coarsely: some use image-plane cues that are ambiguous in depth and rotation, while others use 3D tracks and blobs that lack complete geometry and lose consistency when the viewpoint changes.

4Director is instead conditioned on an explicit 4D scene representation. Each object is reconstructed once from the input image as a canonical mesh, then moved by one prescribed rigid transformation per frame. The authors say this gives an intuitive 3D control interface and stops unobserved geometry from being regenerated independently in every frame.

The controlled scene is rendered as a depth video, which acts as a geometric scaffold. A component called the Motion Adapter transforms that scaffold into video, synthesizing view-consistent appearance, illumination, and non-rigid dynamics.

For training, the authors built RealCOD-Rigid, a new dataset of 20,774 clips annotated with rigid 3D scenes by their automatic pipeline. They also introduce a new metric, Identity-Gated IoU (IG-IoU), which jointly evaluates how well a video follows the prescribed object motion and whether the object's identity is preserved. According to the authors, experiments show that 4Director consistently outperforms prior methods in visual quality and in camera and object control.

Key facts

  • 4Director is a video world model conditioned on an explicit 4D scene representation: each object is reconstructed once from the input image as a canonical mesh and moved by one prescribed rigid transformation per frame.
  • The controlled scene is rendered as a depth video, and a Motion Adapter turns it into video with view-consistent appearance, illumination, and non-rigid dynamics.
  • The authors built RealCOD-Rigid, a training dataset of 20,774 clips annotated with rigid 3D scenes by their automatic pipeline.
  • A new metric, Identity-Gated IoU (IG-IoU), jointly scores adherence to prescribed object motion and preservation of object identity.
  • The authors report that 4Director consistently outperforms prior methods in visual quality and in camera and object control.

Why it matters

The authors say precise control over camera and object motion is essential for professional video production, and that current methods only control objects coarsely. Image-plane cues are ambiguous in depth and rotation; 3D tracks and blobs lack complete geometry and lose consistency across viewpoint changes. 4Director's answer is to give each object a real mesh and an explicit per-frame rigid transformation, so the control signal is geometric rather than a loose hint. The authors also say this stops unobserved geometry from being regenerated independently in every frame.

Who it affects

The framing is professional video production, so the work is aimed at people who need to direct camera and object motion precisely. For researchers, the paper adds a new dataset (RealCOD-Rigid) and a new evaluation metric (IG-IoU) for judging motion adherence together with object identity.

How to use it

The source describes the method, not a product. The workflow it outlines: reconstruct each object once from an input image as a canonical mesh, prescribe one rigid transformation per object per frame, render the scene as a depth video, and let the Motion Adapter turn that into the final video. The source does not mention code, weights or a dataset release.

How solid is it

The claims come from the authors' own summary of their experiments, which they say show consistent gains over prior methods in visual quality and in camera and object control. The source gives no benchmark scores, no IG-IoU values, no margins, and does not name the prior methods compared against. The 20,774-clip dataset size is the one concrete number available.

Risks and caveats

The control interface is built on rigid transformations of meshes, with non-rigid dynamics synthesized by the Motion Adapter rather than directly prescribed. The source does not say how many objects or what kinds of objects the system handles. It also gives no model size, training compute or inference speed, and does not say whether RealCOD-Rigid will be released. Without numbers, the claim of outperforming prior methods cannot be checked from this text.