Atlas
World Labs·
Overview
World Labs' omni world model, pretrained from scratch to operate natively on text, images, video and 3D as a multimodal autoregressive diffusion transformer with every input in one shared spatial context. The headline capability is camera-controlled video, up to one minute at 1440p from one or more reference images, with the camera path supplied as a geometric input rather than described in a prompt. The same model reconstructs scenes into point clouds and 3D Gaussian splats, works with depth maps and camera poses, reframes multi-camera footage and supports parts of a real-to-sim robotics workflow. Early access for selected partners via a request form; no pricing, API or public date announced. World Labs reports that third-party human raters preferred Atlas over rival video models in 75 to 94 percent of pairwise trials depending on the competitor, a vendor-commissioned preference study recorded here as a note, not as a score.
Capabilities and innovations
Benchmarks
No established, comparable metric exists for world models yet. A video model goes from prompt to clip; a world model goes from an image or scene to a navigable space with camera control, which no video or image board measures. No scores and no crown until one exists.
Links
More from World Labs
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.