World Labs launches Atlas, a spatial world model

World Labs, the spatial-intelligence company, introduced Atlas, its next-generation world model. Atlas is an omni model pretrained from scratch to operate natively on text, images, video and 3D. World Labs describes it as a multimodal autoregressive diffusion transformer: every input is combined into one shared spatial context, and the model uses that context to generate what comes next while staying consistent in 3D with everything it has already seen. The company says Atlas's performance improves with more training compute and expects that trend to continue as it keeps scaling.
Atlas spans three broad task families. In camera-controlled generation, it produces images and video from one or more reference images with pixel-perfect camera control, outputting up to 1 minute of video at 1440p resolution; World Labs says this goes beyond coarse text-based camera instructions because Atlas takes precise camera geometry as a native input. In spatial reconstruction, Atlas rebuilds real-world scenes from as few as two or three input images, a figure the company says typically yields faithful results that outperform models built specifically for 3D reconstruction, while the model can also draw on more than a hundred input images for closer recreation of a real environment. In one worked example, Atlas rebuilt Stanford's Main Quad from between two and twenty-five ground-level photos, generating aerial flythroughs the source images never captured. Atlas outputs both 2D image frames and explicit 3D representations, including point clouds and 3D Gaussian splats, the same format used in World Labs's Marble product, which the company says lets the two integrate directly.
The third family, space-time simulation, covers video reframing and robotics. World Labs demonstrated turning footage from as few as three ordinary cell-phone cameras into a bullet-time capture rig, reconstructing a scene from three to five camera views and then letting a user reframe the shot from angles none of the cameras actually held; the company says the footage needed no professional photographers or specialized equipment. For robotics, World Labs captured two environments on cell-phone video using 24 frames each, then used Atlas to simulate different robots navigating the reconstructed spaces and to generate the RGB and depth imagery those robots' onboard cameras would see, with the stated aim of producing varied training and testing data for robotics at scale. Atlas can also generate images and 360-degree panoramas directly from text, though World Labs frames image generation as a secondary capability rather than the model's main focus.
World Labs did not disclose a release date, pricing, API or waitlist access for Atlas, nor did it name individual researchers behind the work, a parameter count, training data size or compute budget, or the specific benchmarks and competing models behind its 'outperforming state of the art' claims.
Key facts
- Atlas is World Labs's next-generation world model, an omni multimodal autoregressive diffusion transformer pretrained from scratch on text, images, video and 3D.
- Camera-controlled generation outputs up to 1 minute of video at 1440p with pixel-perfect camera control from one or more reference images.
- Atlas typically gives faithful spatial reconstructions from as few as two or three input images, and can use over a hundred images for closer recreations; a Stanford Main Quad example used two to twenty-five images.
- For robotics and video reframing, Atlas worked from as few as three ordinary cell-phone cameras or 24 frames per environment to simulate robot navigation or produce bullet-time-style reframed shots.
- World Labs says Atlas will power future versions of Marble and its other products, but gave no release date, pricing, benchmark figures or named researchers.
Why it matters
Atlas is World Labs's attempt at a general-purpose world model: one architecture trained from scratch to generate, reconstruct and simulate 3D worlds rather than a model bolted onto an existing image or video system. Its core idea, a shared spatial context where every image is grounded at a 3D position, borrows the autoregressive sequence approach of language models and combines it with diffusion-based generation and a transformer backbone, letting one model treat generation, reconstruction and simulation as versions of the same task.
Who it affects
World Labs frames Atlas as the engine behind Marble and its other products, so the immediate audience is that product's users and anyone building on it. Beyond that, the model targets creative users who generate scenes and long camera-controlled videos, 3D reconstruction researchers and practitioners, and robotics teams that want to build real-to-sim training and testing environments from ordinary video rather than specialized capture equipment.
How to use it
World Labs says Atlas will power future versions of Marble and other World Labs products, but the announcement gives no release date, pricing, API access or waitlist details for Atlas itself, so there is currently no way to use the model directly outside of what ships inside those products.
How solid is it
The claims come entirely from World Labs's own announcement, not from an independent benchmark or third-party evaluation. The company says Atlas outperforms state-of-the-art specialized 3D reconstruction models and that its performance keeps improving with more training compute, but it names no specific benchmarks, no competing models and discloses no parameter count, training data size or compute budget behind those claims.
Risks and caveats
Every performance claim in the announcement is qualitative rather than backed by published numbers, so the 'outperforming state of the art' language cannot be checked against a specific test or dataset. No individual researchers are credited for the work.