Sep 20, 2024 · 48m · a16z

“The Future of AI is Here” — Fei-Fei Li Unveils the Next Frontier of AI

Fei-Fei Li · 20m spoken Justin Johnson · 17m spoken Martin Casado · 6m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of The a16z Podcast, AI pioneers Dr. Fei-Fei Li and Justin Johnson discuss the evolution of artificial intelligence and unveil their new venture, World Labs. They explain why spatial intelligence—the ability for AI to perceive, reason about, and interact within 3D and 4D environments—represents the next major frontier beyond large language models.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The host as informed peer 3.7 Guest teaching 4.3 Guest disagreement 0.6 The host pushing back 1.4
05100:0015:0030:0045:000:21–3:42 · The host as informed peer 2/10 Disclaimer and a16z Podcast Opening Bumper Martin opens with broad introductory questions inviting Fei-Fei Li and Justin Johnson to share their personal backgrounds in AI. The conversation is welcoming and collaborative without pushback.3:42–6:56 · The host as informed peer 2/10 Fei-Fei Li's Background and the Genesis of ImageNet Fei-Fei Li recounts her transition from physics to computational neuroscience and explains the realization that internet-scale data was the key overlooked driver of model generalization, leading to ImageNet.6:56–9:16 · The host as informed peer 4/10 Breakthrough Epochs: The Role of Compute and AlexNet Martin prompts a discussion on whether breakthroughs come from algorithmic unlocks or compute. Justin illustrates the staggering compute scaling since 2012 by showing AlexNet's training time reduced from six days to under five minutes on modern GPUs.9:16–11:34 · The host as informed peer 6/10 Supervised Learning vs. Unsupervised Data and Implicit Labeling Martin actively challenges the compute-centric narrative by raising data structure and implicit human labeling in CLIP and Transformers. Fei-Fei playfully reframes his argument, noting implicit human labeling holds far more strongly for language than for pixels.11:34–16:58 · The host as informed peer 4/10 The Evolution of Generative AI and Justin Johnson's PhD Milestones Martin asks the guests to trace the shift from predictive computer vision to generative AI. Fei-Fei and Justin map out Justin's PhD milestones spanning image matching, style transfer, and early scene-graph-to-image GAN generation.16:58–20:11 · The host as informed peer 3/10 Founding World Labs and the Pursuit of Spatial Intelligence Martin inquires about the transition from academic research to founding World Labs. Fei-Fei explains her pursuit of visual-spatial intelligence as a fundamental North Star alongside language.20:11–23:29 · The host as informed peer 4/10 Defining Spatial Intelligence and 3D/4D World Representation Martin asks for a precise definition of spatial intelligence and whether it applies to physical or abstract spaces. Justin defines it as perceiving, reasoning, and acting across 3D/4D space-time and details the pivotal impact of NeRF.23:29–32:32 · The host as informed peer 5/10 Contrasting 1D Language Models with 3D Spatial Representation Martin pushes the guests on why multimodal LLMs cannot simply handle spatial tasks. Justin clarifies that LLMs shoehorn data into 1D sequences, whereas true spatial intelligence requires 3D representations front and center.32:32–40:41 · The host as informed peer 4/10 Use Cases for Spatial Intelligence: Interactive Worlds, AR, and Robotics Martin guides the discussion toward concrete commercial applications. Justin and Fei-Fei explore interactive 3D world generation, new media experiences, augmented reality interfaces, and physical robotics.40:41–45:54 · The host as informed peer 4/10 Deep Tech Platform Strategy and Assembling World Labs' Team Martin asks how World Labs balances deep tech platform ambitions with specific vertical applications. Fei-Fei outlines their foundational deep tech positioning and describes recruiting co-founders Ben Mildenhall and Christoph Lassner.45:54–47:34 · The host as informed peer 3/10 Reaching the North Star and the Expanding Frontier of AI Martin asks how the team will evaluate reaching their ultimate goal. Fei-Fei highlights real-world deployment impact, while Justin emphasizes that understanding a 4D universe is an expanding, infinite journey.0:21–3:42 · Guest teaching 2/10 Disclaimer and a16z Podcast Opening Bumper Martin opens with broad introductory questions inviting Fei-Fei Li and Justin Johnson to share their personal backgrounds in AI. The conversation is welcoming and collaborative without pushback.3:42–6:56 · Guest teaching 4/10 Fei-Fei Li's Background and the Genesis of ImageNet Fei-Fei Li recounts her transition from physics to computational neuroscience and explains the realization that internet-scale data was the key overlooked driver of model generalization, leading to ImageNet.6:56–9:16 · Guest teaching 6/10 Breakthrough Epochs: The Role of Compute and AlexNet Martin prompts a discussion on whether breakthroughs come from algorithmic unlocks or compute. Justin illustrates the staggering compute scaling since 2012 by showing AlexNet's training time reduced from six days to under five minutes on modern GPUs.9:16–11:34 · Guest teaching 5/10 Supervised Learning vs. Unsupervised Data and Implicit Labeling Martin actively challenges the compute-centric narrative by raising data structure and implicit human labeling in CLIP and Transformers. Fei-Fei playfully reframes his argument, noting implicit human labeling holds far more strongly for language than for pixels.11:34–16:58 · Guest teaching 5/10 The Evolution of Generative AI and Justin Johnson's PhD Milestones Martin asks the guests to trace the shift from predictive computer vision to generative AI. Fei-Fei and Justin map out Justin's PhD milestones spanning image matching, style transfer, and early scene-graph-to-image GAN generation.16:58–20:11 · Guest teaching 4/10 Founding World Labs and the Pursuit of Spatial Intelligence Martin inquires about the transition from academic research to founding World Labs. Fei-Fei explains her pursuit of visual-spatial intelligence as a fundamental North Star alongside language.20:11–23:29 · Guest teaching 5/10 Defining Spatial Intelligence and 3D/4D World Representation Martin asks for a precise definition of spatial intelligence and whether it applies to physical or abstract spaces. Justin defines it as perceiving, reasoning, and acting across 3D/4D space-time and details the pivotal impact of NeRF.23:29–32:32 · Guest teaching 6/10 Contrasting 1D Language Models with 3D Spatial Representation Martin pushes the guests on why multimodal LLMs cannot simply handle spatial tasks. Justin clarifies that LLMs shoehorn data into 1D sequences, whereas true spatial intelligence requires 3D representations front and center.32:32–40:41 · Guest teaching 4/10 Use Cases for Spatial Intelligence: Interactive Worlds, AR, and Robotics Martin guides the discussion toward concrete commercial applications. Justin and Fei-Fei explore interactive 3D world generation, new media experiences, augmented reality interfaces, and physical robotics.40:41–45:54 · Guest teaching 3/10 Deep Tech Platform Strategy and Assembling World Labs' Team Martin asks how World Labs balances deep tech platform ambitions with specific vertical applications. Fei-Fei outlines their foundational deep tech positioning and describes recruiting co-founders Ben Mildenhall and Christoph Lassner.45:54–47:34 · Guest teaching 3/10 Reaching the North Star and the Expanding Frontier of AI Martin asks how the team will evaluate reaching their ultimate goal. Fei-Fei highlights real-world deployment impact, while Justin emphasizes that understanding a 4D universe is an expanding, infinite journey.0:21–3:42 · Guest disagreement 0/10 Disclaimer and a16z Podcast Opening Bumper Martin opens with broad introductory questions inviting Fei-Fei Li and Justin Johnson to share their personal backgrounds in AI. The conversation is welcoming and collaborative without pushback.3:42–6:56 · Guest disagreement 0/10 Fei-Fei Li's Background and the Genesis of ImageNet Fei-Fei Li recounts her transition from physics to computational neuroscience and explains the realization that internet-scale data was the key overlooked driver of model generalization, leading to ImageNet.6:56–9:16 · Guest disagreement 1/10 Breakthrough Epochs: The Role of Compute and AlexNet Martin prompts a discussion on whether breakthroughs come from algorithmic unlocks or compute. Justin illustrates the staggering compute scaling since 2012 by showing AlexNet's training time reduced from six days to under five minutes on modern GPUs.9:16–11:34 · Guest disagreement 3/10 Supervised Learning vs. Unsupervised Data and Implicit Labeling Martin actively challenges the compute-centric narrative by raising data structure and implicit human labeling in CLIP and Transformers. Fei-Fei playfully reframes his argument, noting implicit human labeling holds far more strongly for language than for pixels.11:34–16:58 · Guest disagreement 1/10 The Evolution of Generative AI and Justin Johnson's PhD Milestones Martin asks the guests to trace the shift from predictive computer vision to generative AI. Fei-Fei and Justin map out Justin's PhD milestones spanning image matching, style transfer, and early scene-graph-to-image GAN generation.16:58–20:11 · Guest disagreement 0/10 Founding World Labs and the Pursuit of Spatial Intelligence Martin inquires about the transition from academic research to founding World Labs. Fei-Fei explains her pursuit of visual-spatial intelligence as a fundamental North Star alongside language.20:11–23:29 · Guest disagreement 0/10 Defining Spatial Intelligence and 3D/4D World Representation Martin asks for a precise definition of spatial intelligence and whether it applies to physical or abstract spaces. Justin defines it as perceiving, reasoning, and acting across 3D/4D space-time and details the pivotal impact of NeRF.23:29–32:32 · Guest disagreement 2/10 Contrasting 1D Language Models with 3D Spatial Representation Martin pushes the guests on why multimodal LLMs cannot simply handle spatial tasks. Justin clarifies that LLMs shoehorn data into 1D sequences, whereas true spatial intelligence requires 3D representations front and center.32:32–40:41 · Guest disagreement 0/10 Use Cases for Spatial Intelligence: Interactive Worlds, AR, and Robotics Martin guides the discussion toward concrete commercial applications. Justin and Fei-Fei explore interactive 3D world generation, new media experiences, augmented reality interfaces, and physical robotics.40:41–45:54 · Guest disagreement 0/10 Deep Tech Platform Strategy and Assembling World Labs' Team Martin asks how World Labs balances deep tech platform ambitions with specific vertical applications. Fei-Fei outlines their foundational deep tech positioning and describes recruiting co-founders Ben Mildenhall and Christoph Lassner.45:54–47:34 · Guest disagreement 0/10 Reaching the North Star and the Expanding Frontier of AI Martin asks how the team will evaluate reaching their ultimate goal. Fei-Fei highlights real-world deployment impact, while Justin emphasizes that understanding a 4D universe is an expanding, infinite journey.0:21–3:42 · The host pushing back 0/10 Disclaimer and a16z Podcast Opening Bumper Martin opens with broad introductory questions inviting Fei-Fei Li and Justin Johnson to share their personal backgrounds in AI. The conversation is welcoming and collaborative without pushback.3:42–6:56 · The host pushing back 0/10 Fei-Fei Li's Background and the Genesis of ImageNet Fei-Fei Li recounts her transition from physics to computational neuroscience and explains the realization that internet-scale data was the key overlooked driver of model generalization, leading to ImageNet.6:56–9:16 · The host pushing back 2/10 Breakthrough Epochs: The Role of Compute and AlexNet Martin prompts a discussion on whether breakthroughs come from algorithmic unlocks or compute. Justin illustrates the staggering compute scaling since 2012 by showing AlexNet's training time reduced from six days to under five minutes on modern GPUs.9:16–11:34 · The host pushing back 6/10 Supervised Learning vs. Unsupervised Data and Implicit Labeling Martin actively challenges the compute-centric narrative by raising data structure and implicit human labeling in CLIP and Transformers. Fei-Fei playfully reframes his argument, noting implicit human labeling holds far more strongly for language than for pixels.11:34–16:58 · The host pushing back 1/10 The Evolution of Generative AI and Justin Johnson's PhD Milestones Martin asks the guests to trace the shift from predictive computer vision to generative AI. Fei-Fei and Justin map out Justin's PhD milestones spanning image matching, style transfer, and early scene-graph-to-image GAN generation.16:58–20:11 · The host pushing back 0/10 Founding World Labs and the Pursuit of Spatial Intelligence Martin inquires about the transition from academic research to founding World Labs. Fei-Fei explains her pursuit of visual-spatial intelligence as a fundamental North Star alongside language.20:11–23:29 · The host pushing back 2/10 Defining Spatial Intelligence and 3D/4D World Representation Martin asks for a precise definition of spatial intelligence and whether it applies to physical or abstract spaces. Justin defines it as perceiving, reasoning, and acting across 3D/4D space-time and details the pivotal impact of NeRF.23:29–32:32 · The host pushing back 3/10 Contrasting 1D Language Models with 3D Spatial Representation Martin pushes the guests on why multimodal LLMs cannot simply handle spatial tasks. Justin clarifies that LLMs shoehorn data into 1D sequences, whereas true spatial intelligence requires 3D representations front and center.32:32–40:41 · The host pushing back 1/10 Use Cases for Spatial Intelligence: Interactive Worlds, AR, and Robotics Martin guides the discussion toward concrete commercial applications. Justin and Fei-Fei explore interactive 3D world generation, new media experiences, augmented reality interfaces, and physical robotics.40:41–45:54 · The host pushing back 1/10 Deep Tech Platform Strategy and Assembling World Labs' Team Martin asks how World Labs balances deep tech platform ambitions with specific vertical applications. Fei-Fei outlines their foundational deep tech positioning and describes recruiting co-founders Ben Mildenhall and Christoph Lassner.45:54–47:34 · The host pushing back 0/10 Reaching the North Star and the Expanding Frontier of AI Martin asks how the team will evaluate reaching their ultimate goal. Fei-Fei highlights real-world deployment impact, while Justin emphasizes that understanding a 4D universe is an expanding, infinite journey.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 0% · guest 100%0:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%33:00 · the host 0% · guest 100%33:00 · the host 0% · guest 100%36:00 · the host 0% · guest 100%36:00 · the host 0% · guest 100%39:00 · the host 0% · guest 100%39:00 · the host 0% · guest 100%42:00 · the host 0% · guest 100%42:00 · the host 0% · guest 100%45:00 · the host 0% · guest 100%45:00 · the host 0% · guest 100%48:00 · the host 0% · guest 0%48:00 · the host 0% · guest 0%
Sharpest disagreement ▶ 10:46 Fei-Fei Li playful rejection of Martin's premise on human data

When Martin asserts that self-attention models rely on implicit human labeling, Fei-Fei playfully responds that she knew he would say that, directly pushing back that his point applies far more to language than to pixels.

Hardest push from the host ▶ 10:38 Martin Casado challenges the compute-only narrative using data labeling arguments

Martin positions himself as a naive listener to directly challenge the compute-driven 'bitter lesson' narrative, arguing forcefully that human-labeled structure in CLIP alt-tags and text data is what truly unlocks deep learning.

Biggest teaching moment ▶ 27:08 Justin Johnson explains the fundamental 1D limitation of multimodal LLMs

Justin educates the host on model architectures, explaining that LLMs operate on 1D token sequences and shoehorn visual inputs, whereas true spatial intelligence requires 3D representations built into the model's core.

The host holds their own ▶ 9:16 Martin Casado details data-centric unlocks against compute-centric assumptions

Martin demonstrates deep familiarity with AI history by contrasting the 'bitter lesson' compute argument with specific data-centric counterexamples like ImageNet, sentence structure, and CLIP alt-tags.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
Disclaimer and a16z Podcast Opening Bumper 2200 Martin opens with broad introductory questions inviting Fei-Fei Li and Justin Johnson to share their personal backgrounds in AI. The conversation is welcoming and collaborative without pushback.
Fei-Fei Li's Background and the Genesis of ImageNet 2400 Fei-Fei Li recounts her transition from physics to computational neuroscience and explains the realization that internet-scale data was the key overlooked driver of model generalization, leading to ImageNet.
Breakthrough Epochs: The Role of Compute and AlexNet 4612 Martin prompts a discussion on whether breakthroughs come from algorithmic unlocks or compute. Justin illustrates the staggering compute scaling since 2012 by showing AlexNet's training time reduced from six days to under five minutes on modern GPUs.
Supervised Learning vs. Unsupervised Data and Implicit Labeling 6536 Martin actively challenges the compute-centric narrative by raising data structure and implicit human labeling in CLIP and Transformers. Fei-Fei playfully reframes his argument, noting implicit human labeling holds far more strongly for language than for pixels.
The Evolution of Generative AI and Justin Johnson's PhD Milestones 4511 Martin asks the guests to trace the shift from predictive computer vision to generative AI. Fei-Fei and Justin map out Justin's PhD milestones spanning image matching, style transfer, and early scene-graph-to-image GAN generation.
Founding World Labs and the Pursuit of Spatial Intelligence 3400 Martin inquires about the transition from academic research to founding World Labs. Fei-Fei explains her pursuit of visual-spatial intelligence as a fundamental North Star alongside language.
Defining Spatial Intelligence and 3D/4D World Representation 4502 Martin asks for a precise definition of spatial intelligence and whether it applies to physical or abstract spaces. Justin defines it as perceiving, reasoning, and acting across 3D/4D space-time and details the pivotal impact of NeRF.
Contrasting 1D Language Models with 3D Spatial Representation 5623 Martin pushes the guests on why multimodal LLMs cannot simply handle spatial tasks. Justin clarifies that LLMs shoehorn data into 1D sequences, whereas true spatial intelligence requires 3D representations front and center.
Use Cases for Spatial Intelligence: Interactive Worlds, AR, and Robotics 4401 Martin guides the discussion toward concrete commercial applications. Justin and Fei-Fei explore interactive 3D world generation, new media experiences, augmented reality interfaces, and physical robotics.
Deep Tech Platform Strategy and Assembling World Labs' Team 4301 Martin asks how World Labs balances deep tech platform ambitions with specific vertical applications. Fei-Fei outlines their foundational deep tech positioning and describes recruiting co-founders Ben Mildenhall and Christoph Lassner.
Reaching the North Star and the Expanding Frontier of AI 3300 Martin asks how the team will evaluate reaching their ultimate goal. Fei-Fei highlights real-world deployment impact, while Justin emphasizes that understanding a 4D universe is an expanding, infinite journey.

Statements from this episode (15)

Opinion
Fei-Fei Li: Spatial intelligence is as fundamental to AI as language
“Visual spatial intelligence is so fundamental. It's as fundamental as language.”
Fei-Fei Li Sep 20, 2024 ▶ 18:58
Opinion
Fei-Fei Li: AI is undergoing a Cambrian explosion beyond text
“And now I think we're in the middle of a Cambrian explosion. In almost a literal sense, because now in addition to texts, you're seeing pixels, videos, audios, all coming out with possible AI applications and a model.”
Fei-Fei Li Sep 20, 2024 ▶ 1:20
Assertion Supported
Justin Johnson: AlexNet trained for six days on two consumer GPUs
“That AlexNet was a sixty million parameter deep neural network and it was trained for six days on two GTX five eighties, which was the top consumer card at the time, which came out in 2010.”
Justin Johnson Sep 20, 2024 ▶ 7:56
Assertion Supported
Justin Johnson: AlexNet's 6-day training takes under 5 minutes on GB200
“So I ran the numbers last night, like that two week training run, that of six days on two GTX five eighties, if you scale, it comes out to just under five minutes on a single GB 200.”
Justin Johnson Sep 20, 2024 ▶ 8:26
Assertion Not checkable as stated
Fei-Fei Li: AlexNet's architecture mirrored 1980s models, unlocked by GPUs
“The 2012 AlexNet paper on ImageNet Challenge is literally a very classic model, and that is the convolutional neural network model, and that was published in 19 eighties, the first paper. I remember as a graduate student learning that, and it more or less also…”
Fei-Fei Li Sep 20, 2024 ▶ 8:41
Insight
Fei-Fei Li: Language data contains more implicit human labeling than pixels
“Yes, philosophically, that's a really important question. But that actually is more true in language than pixels.”
Fei-Fei Li Sep 20, 2024 ▶ 10:49
Insight
Justin Johnson: AI is shifting from analyzing web data to sensor data
“The previous decade had mostly been about understanding data that already exists. But the next decade was going to be about understanding new data.”
Justin Johnson Sep 20, 2024 ▶ 21:32
Assertion Supported
Justin Johnson: Ben Mildenhall's 2020 NeRF paper ignited 3D computer vision
“In 2020, you asked about breakthrough moments. There was a really big breakthrough moment from our co-founder Ben Mildenhall at the time with his paper, NERF neural radiance fields. And that was a very simple, very clear way of backing out three D structure fr…”
Justin Johnson Sep 20, 2024 ▶ 22:54
Insight
Justin Johnson: Multimodal LLMs shoehorn visual data into 1D token sequences
“And now the multimodal LLMs that we're seeing now, you kind of end up shoehorning the other modalities into this underlying representation of a one D sequence of tokens. Now when we move to spatial intelligence, it's kind of going the other way. Where we're sa…”
Justin Johnson Sep 20, 2024 ▶ 27:44
Insight
Fei-Fei Li: The 3D world follows physics, unlike human-generated language
“Language is fundamentally a purely generated signal. There's no language out there. You don't go out in the nature and there's words written in the sky for you. Whatever data you feed in, you pretty much can just Somehow regurgitate with enough generalizabilit…”
Fei-Fei Li Sep 20, 2024 ▶ 28:33
Prediction Not checkable as stated
Justin Johnson: Native 3D AI representations will outperform 2D video generation
“Modeling the two D projections of a dynamic three D world is, is a function that probably can be modeled, but by putting a three D representation into the heart of a model, there's just going to be a better fit between the kind of representation that the model…”
Justin Johnson Sep 20, 2024 ▶ 31:00
Assertion Not checkable as stated
Justin Johnson: Creating interactive 3D worlds currently costs hundreds of millions
“Because we already have the ability to create virtual interactive worlds but it costs hundreds and hundreds of millions of dollars and a ton of development time.”
Justin Johnson Sep 20, 2024 ▶ 33:53
Prediction Not checkable as stated
Fei-Fei Li: Spatial intelligence will be the operating system for AR/VR
“Suddenly this piece of technology is, is going to be the operating system, basically for AR, VR MixR.”
Fei-Fei Li Sep 20, 2024 ▶ 38:39
Prediction Not checkable as stated
Justin Johnson: Seamless mixed reality will deprecate phones, TVs, and monitors
“If you've got the ability to seamlessly blend virtual content with the physical world, it kind of deprecates the need for all of those.”
Justin Johnson Sep 20, 2024 ▶ 39:35
Opinion
Justin Johnson: VR headsets like Vision Pro lack mass market readiness
“But I think the reality is it's just not there yet as a platform for mass market appeal.”
Justin Johnson Sep 20, 2024 ▶ 41:45
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.