World Labs describes Atlas as an omni-model trained from scratch on text, images, video, and 3D data. Every input gets anchored to a specific position in 3D space rather than processed as a flat sequence. The company calls this shared spatial understanding "spatial context," and it's what the model uses to generate each new frame or viewpoint. According to World Labs, this anchoring separates Atlas from pure language or video models.
Fei-Fei Li laid out this exact problem in a November 2025 essay. Current multimodal language models and video diffusion models break data into one- or two-dimensional sequences, she argued, which makes even simple spatial tasks needlessly hard. What's needed are architectures that organize tokenization, context, and memory in a 3D- or 4D-aware way.
One minute of video at 1440p
For camera-controlled generation, Atlas takes one or more images and produces new views at freely chosen camera positions and angles. Camera movement is passed as a direct geometric input rather than described through text prompts, as many video models require.
The model outputs up to one minute of video at 1440p. Users can control every shot themselves instead of "pulling the lever on a slot machine," as World Labs put it, drawing a line between controlled generation and random output.






