World Labs Unveils Atlas to Unify 3D and Video Generation
Spatial intelligence startup World Labs has unveiled Atlas, a foundation model that merges 3D reconstruction and video generation to simplify virtual world-building and robotics simulation.

World Labs, the spatial intelligence startup cofounded by Fei-Fei Li, has introduced Atlas, an omni foundation model that integrates text, images, video, and 3D data into a single architecture. Operating as a multimodal autoregressive diffusion transformer, Atlas processes native camera pose inputs alongside visual data, anchoring each image to a specific 3D position. The system can generate up to one minute of 1440p video from reference images or reconstruct detailed 3D scenes, outputting point clouds and Gaussian splats from as few as two or three photos.
In human evaluation trials for camera-controlled generation, raters preferred Atlas over rival models by significant margins: 75 percent over MiniMax H3, 81 percent over Gemini Omni Flash, 86 percent over Happy Horse 1.1, 93 percent over FLUX 3, and 94 percent over Seedance 2.5. However, World Labs noted that competitors received camera paths via text descriptions rather than native pose inputs. On 3D reconstruction benchmarks, Atlas achieved lower absolute-relative pointmap error than specialized models like VGGT, Depth Anything 3, MapAnything, and Pi3X across the DTU, ETH3D, KITTI, ScanNet, and Tanks and Temples datasets.
For developers and researchers, Atlas represents a shift away from fragmented pipelines that rely on separate depth estimators, video diffusion models, and NeRF or Gaussian splat tools. By feeding camera geometry directly into the model as numeric coordinates, practitioners gain precise camera control without relying on imprecise text prompts. This unified approach simplifies real-to-sim workflows for robotics, allowing developers to reconstruct entire environments from roughly 24 frames of mobile phone video to generate diverse training data.
Currently, Atlas is only accessible through an early access request form, with no public API, pricing, model card, or code released yet. World Labs plans to integrate the model into Marble, its platform for building explorable 3D worlds.
This is our own summary of reporting by AlphaSignal



