Reconstructing a 3D scene used to require hundreds of carefully captured photographs, expensive equipment, and a healthy dose of patience. World Labs, the spatial intelligence company co-founded by Stanford AI luminary Fei-Fei Li, just compressed that workflow down to two or three snapshots.

The company’s new model, called Atlas, is a pretrained multimodal autoregressive diffusion transformer. In plainer terms: it’s an AI system that understands text, images, video, and 3D data all at once, and can use that understanding to fill in everything a camera didn’t capture. The result is full 3D scene reconstruction, novel view synthesis, and camera-controlled video generation, all from minimal input.

What Atlas actually does

Atlas operates as what World Labs calls an “omni world model” for spatial intelligence. Feed it between one and six reference images, and it can generate up to one minute of video at 1440p resolution while maintaining precise camera control along exact paths you specify.

The 3D reconstruction capability is where things get genuinely impressive. Traditional photogrammetry pipelines typically demand anywhere from 100 to 300+ images shot from carefully planned angles to build a usable 3D model. Atlas collapses that requirement to as few as two or three input images, producing explicit 3D representations like point clouds and Gaussian splats with what the company describes as remarkable fidelity.