Jun 4, 2025 · 22m · a16z
How Fei-Fei Li Is Rebuilding AI for the Real World
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of The a16z Podcast, AI pioneer Dr. Fei-Fei Li and General Partner Martin Casado discuss why artificial intelligence must evolve beyond language models to master 3D spatial intelligence. They detail the founding mission, scientific underpinnings, and real-world applications of World Labs in developing Large World Models.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the host, purple is the guest (3 minute bins)
In a highly amicable interview, the closest moment to friction occurs when Fei-Fei interrupts Martin to correct his evolutionary timeline, clarifying that spatial vision dates back 500 million years to trilobites rather than 4 million.
Hardest push from the host ▶ 6:01 Challenging the need beyond LLMsThe host directly asks why LLMs are not enough and questions the rationale for building another foundation model company when language models are dominating.
Biggest teaching moment ▶ 6:05 Reframing language vs spatial intelligenceFei-Fei reframes the entire AI paradigm by explaining that language is merely a recent, lossy human construct, whereas 3D spatial perception is the true foundation of embodied intelligence.
The host holds their own ▶ 16:51 Probing 2D limitationsThe host tests the technical premise of world models by asking why existing 2D video outputs cannot simply be used instead of dedicated 3D representations.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The host as informed peer | Guest teaching | Guest disagreement | The host pushing back | Why |
|---|---|---|---|---|---|---|
| Introducing Fei-Fei Li: The Godmother of AI | 1 | 1 | 0 | 0 | The host provides friendly setup questions introducing guest Fei-Fei Li and investor Martin Casado. The dynamic is purely collaborative and promotional as they share how World Labs was conceived. | |
| Reflecting on a Decade of AI Progress | 2 | 4 | 1 | 2 | The host asks why LLMs are not enough and why a new foundation model company is necessary. Fei-Fei educates on how language is a recent, lossy encoding of thought while spatial 3D intelligence is fundamental to living things. | |
| Evolutionary Perspective on Language vs. Spatial Intelligence | 1 | 4 | 1 | 0 | The guests outline the evolutionary timeline of vision versus language models. Fei-Fei gently corrects Martin's evolutionary estimate by pointing out spatial intelligence dates back 500 million years to trilobites. | |
| Spatial Intelligence as the Engine of Human Breakthroughs | 1 | 2 | 0 | 0 | The host prompts for concrete applications of world models. Fei-Fei and Martin walk through practical use cases in robotics, design, and generative multiverses. | |
| Why 3D Representation Is Essential: The Tree Analogy and Monocular Vision Story | 2 | 4 | 1 | 1 | The host questions why 2D representations cannot suffice. Fei-Fei illustrates the essential nature of 3D vision using an anecdote about losing her stereo vision after an eye injury and feeling terrified to drive. | |
| The Scientific Foundations and Interdisciplinary Team of World Labs | 1 | 3 | 0 | 0 | The host asks about the state of research in 3D AI relative to LLMs. Fei-Fei details the academic contributions of her co-founders across NeRFs, Gaussian splatting, and diffusion models. |