World Labs has launched Atlas, described as the world's first multimodal world model capable of generating image and video frames with pixel-perfect camera control while simultaneously reconstructing them in 3D. Built on a unified multimodal autoregressive diffusion transformer pretrained from scratch, Atlas can produce 1-minute 1440p videos from as few as seven reference images, outperform leading open-source 3D reconstruction models, and generate photorealistic RGB and depth sensor data for robotics simulation from casual photos. Early access sign-ups are opening in the coming weeks via worldlabs.ai. Atlas represents a significant convergence of generative AI, 3D vision, and robotics simulation into a single scalable foundation model architecture.