
Sep 4, 2026 · 45 min
Atlas turns sparse views into navigable worlds
Fei Fei Li: The Race to Build World Models For AI
The conversation examines whether new-view prediction can become a foundational capability for spatial intelligence, from creative tools to robotics.
- 1Atlas predicts unseen viewpoints while combining generation, reconstruction, spatial conditioning, and simulation.
- 2World Labs reduces complex scene capture to a few images by using learned structure to fill missing views.
- 3The model’s next challenges are richer dynamics, reliable editability, and connections between real-world observation and robotic action.
Don't miss
Fei-Fei Li recounts how a smaller model convincingly reconstructed a familiar NeRF scene, giving the team confidence that Atlas was viable.
The brief
Fei-Fei Li, Justin Johnson, and Ben Mildenhall introduce Atlas, World Labs’ world model for generating, reconstructing, and simulating scenes from sparse observations.
Its core idea is new-view prediction: rather than predicting the next token or frame, Atlas generates what a scene should look like from unseen positions in space and time.
The team connects Atlas to the Matrix’s bullet-time effect, where synthesized viewpoints could replace the hundreds of cameras traditionally needed for a moving shot.
Atlas combines camera parameters, RGB imagery, depth, generation, and reconstruction to create spatially consistent fly-throughs instead of merely plausible individual images.
A smaller model’s convincing reconstruction of a familiar NeRF scene became the breakthrough moment that persuaded the team Atlas could work.
The founders see a path from creative tools and architecture to robotics, while acknowledging that dynamics, editability, and real-world data remain difficult.
Listen to the full episode and explore every guest, topic, and moment on PodLume.

Fei-Fei Li
Justin Johnson
Ben Mildenhall
World Labs