Sep 4, 2026 · 43m · a16z
Why Fei-Fei Li Is Betting on Spatial Intelligence
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of The a16z Show, World Labs co-founders Fei-Fei Li, Justin Johnson, and Ben Mildenhall join host Martin Casado to unveil Atlas, their pioneering world model that unifies 3D reconstruction and generative visual AI through new view prediction. They explore how grounding pixels in spatial geometry revolutionizes creative production, powers real-to-sim robotics pipelines, and serves as an evolutionary prerequisite for artificial general intelligence.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the host, purple is the guest (3 minute bins)
Justin Johnson bluntly interrupts Casado's colloquial LLM explanation with 'No, no' and grounds the definition in formal complexity theory.
Hardest push from the host ▶ 36:45 Challenging dynamics versus reconstructionCasado relays tough external critique about static models and directly questions whether reconstruction and dynamics are inherently at odds.
Biggest teaching moment ▶ 41:01 Complexity theory tutorialJohnson corrects Casado's informal conception of AI completeness by explaining its formal linkage to Turing completeness and 3-SAT reduction.
The host holds their own ▶ 42:20 Casado builds on the AI complete thought experimentCasado immediately demonstrates conceptual grasp of Johnson's premise by extending the visual AI completeness example to proof-solving and whiteboard reveals.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The host as informed peer | Guest teaching | Guest disagreement | The host pushing back | Why |
|---|---|---|---|---|---|---|
| Introducing Atlas: Core Capabilities and Bullet-Time Demos | 3 | 2 | 1 | 1 | Casado sets the stage with open, foundational questions about Atlas, asking how it fundamentally differs from standard video models. Johnson and Mildenhall explain new view prediction and spatial context collaboratively without friction. | |
| Joint Reconstruction and Generation in a Unified Architecture | 2 | 4 | 1 | 2 | Casado asks whether Atlas is simply scaled-up video tech or a novel architecture and openly admits he is unfamiliar with 3D representations. Johnson and Li explain depth maps, camera poses, and the historical bifurcation between generation and reconstruction. | |
| Defining Spatial Intelligence and Grounded Geometry | 3 | 3 | 0 | 1 | Casado asks the guests to define spatial intelligence and explain why viewpoint prediction represents a critical milestone. Li details the progression from pixels to geometry, physics, and downstream action planning. | |
| The Marble Predecessor and Establishing Spatial Scaling Laws | 3 | 3 | 0 | 1 | Casado probes why World Labs released Marble before Atlas, prompting Johnson to detail the constraints of Gaussian splats as an output bottleneck. Li notes that prior to their work, no scaling law existed for spatial intelligence. | |
| Sparse vs Dense Reconstruction and Visual Context Scaling | 3 | 4 | 0 | 1 | Casado invites Mildenhall to contrast dense and sparse reconstruction, asking simple questions to clarify why dense capture is so arduous. Mildenhall and Johnson explain the marriage of triangulation with generative hole-filling across context windows. | |
| Conviction in Scaling and the Garden Table Breakthrough | 3 | 4 | 2 | 2 | Casado asks whether the team was truly confident and questions if the architecture is nearing the end of its scaling curve. Li firmly rejects this premise, stating they are only at the beginning, before recounting the breakthrough garden table capture. | |
| Creative and Industrial Use Cases: Statefulness and Architecture | 3 | 2 | 0 | 0 | Casado prompts Mildenhall to explain how Atlas serves creative industries, architecture, and design. Mildenhall walks through traditional 3D design bottlenecks and the necessity of persistent, stateful environments. | |
| Robotics, Real-to-Sim Data Generation, and Neural Simulators | 3 | 4 | 1 | 2 | Casado points out that the connection between a visual world model and robotics is not obvious. Li and Johnson walk him through real-to-sim pipelines, domain randomization, and using learned neural simulators directly as action planners. | |
| Temporal Dynamics, 4D Video, and Model Editability | 4 | 4 | 2 | 4 | Casado relays external expert criticism that the model lacks dynamics and pushes on whether motion conflicts with 3D reconstruction accuracy. Johnson explains that exposure to dynamics during pre-training improves static reconstructions, while Li clarifies control as editability. | |
| AI Completeness and the Evolutionary Mandate of Movement | 4 | 5 | 3 | 1 | Johnson introduces AI completeness, and when Casado offers a colloquial LLM definition, Johnson directly corrects him using Turing completeness and NP reduction to 3-SAT. Casado quickly adapts the premise to visual framing, and Li connects new viewpoints to evolutionary biology. |