Nov 25, 2025 · 1h 0m · latent-space

After LLMs: Spatial Intelligence and World Models — Fei-Fei Li & Justin Johnson, World Labs

Justin Johnson · 22m spoken Fei-Fei Li · 20m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

AI pioneers Fei-Fei Li and Justin Johnson discuss the founding of World Labs and their flagship 3D generative platform, Marble, exploring how the frontier of artificial intelligence is moving beyond text-based large language models toward spatial intelligence, physical world models, and interactive 3D environments.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 4.7 Guest teaching 4.4 Guest disagreement 2.1 The hosts pushing back 1.9
05100:0015:0030:0045:001:00:001:12–3:30 · The hosts as informed peer 3/10 Origins of World Labs and Founder Backgrounds The hosts open with light background questions asking how Fei-Fei Li and Justin Johnson started World Labs. The dynamic is purely conversational, warm, and collegial with no friction or substantive debate.3:31–13:35 · The hosts as informed peer 5/10 Compute Scaling, Academic Research, and Future Hardware Alessio and Swix probe whether academic open challenges like ImageNet still work under modern compute and commercial pressures. Fei-Fei clarifies that the issue is academic under-resourcing rather than open versus closed research, while Justin pushes back slightly against the 'hardware lottery' assumption by analyzing GPU performance per watt.13:36–20:41 · The hosts as informed peer 4/10 The Evolution of Image Captioning and Dense Captioning The guests recount the historical lineage of neural image captioning and dense captioning from their early Stanford lab days. The hosts act primarily as engaged facilitators asking technical follow-ups about forward passes and real-time execution.20:42–29:56 · The hosts as informed peer 6/10 Pixel Maximalism, Physical Laws, and Machine Understanding Alessio brings up a technical paper on LLMs failing to represent orbital force vectors, sparking a debate on whether latent models can truly learn causal physical laws versus pattern matching. Fei-Fei and Justin draw clear philosophical distinctions between statistical pattern fitting and true human-style understanding.29:56–42:11 · The hosts as informed peer 5/10 Marble Architecture, Gaussian Splats, and Practical Applications The hosts inquire into the fundamental data representations underlying Marble and Gaussian splats, asking why embodied robotics wasn't emphasized. Fei-Fei clarifies that simulation for robotics is already featured prominently in their plans, explaining the bridge between synthetic data generation and embodied learning.42:12–55:50 · The hosts as informed peer 6/10 Defining Spatial Intelligence, Embodiment, and Theory Building Fei-Fei immediately rejects the common industry framing of 'a data center full of Einsteins', reframing intelligence as multimodal and highlighting how spatial perception took 540 million years of biological evolution. Alessio counters by noting how language formalizes physical laws like Newton's, leading Justin to separate embodied experience from symbolic theory building.55:50–58:05 · The hosts as informed peer 4/10 Re-evaluating Model Architectures: Transformers as Set Processors When Swix asks if sequence-to-sequence modeling and attention are obsolete for world models, Justin directly corrects the technical premise by explaining that transformers are natively permutation-equivariant set processors rather than sequence models. Swix accepts the clarification as Justin walks through positional embeddings and token-level operations.1:12–3:30 · Guest teaching 1/10 Origins of World Labs and Founder Backgrounds The hosts open with light background questions asking how Fei-Fei Li and Justin Johnson started World Labs. The dynamic is purely conversational, warm, and collegial with no friction or substantive debate.3:31–13:35 · Guest teaching 4/10 Compute Scaling, Academic Research, and Future Hardware Alessio and Swix probe whether academic open challenges like ImageNet still work under modern compute and commercial pressures. Fei-Fei clarifies that the issue is academic under-resourcing rather than open versus closed research, while Justin pushes back slightly against the 'hardware lottery' assumption by analyzing GPU performance per watt.13:36–20:41 · Guest teaching 3/10 The Evolution of Image Captioning and Dense Captioning The guests recount the historical lineage of neural image captioning and dense captioning from their early Stanford lab days. The hosts act primarily as engaged facilitators asking technical follow-ups about forward passes and real-time execution.20:42–29:56 · Guest teaching 5/10 Pixel Maximalism, Physical Laws, and Machine Understanding Alessio brings up a technical paper on LLMs failing to represent orbital force vectors, sparking a debate on whether latent models can truly learn causal physical laws versus pattern matching. Fei-Fei and Justin draw clear philosophical distinctions between statistical pattern fitting and true human-style understanding.29:56–42:11 · Guest teaching 4/10 Marble Architecture, Gaussian Splats, and Practical Applications The hosts inquire into the fundamental data representations underlying Marble and Gaussian splats, asking why embodied robotics wasn't emphasized. Fei-Fei clarifies that simulation for robotics is already featured prominently in their plans, explaining the bridge between synthetic data generation and embodied learning.42:12–55:50 · Guest teaching 6/10 Defining Spatial Intelligence, Embodiment, and Theory Building Fei-Fei immediately rejects the common industry framing of 'a data center full of Einsteins', reframing intelligence as multimodal and highlighting how spatial perception took 540 million years of biological evolution. Alessio counters by noting how language formalizes physical laws like Newton's, leading Justin to separate embodied experience from symbolic theory building.55:50–58:05 · Guest teaching 8/10 Re-evaluating Model Architectures: Transformers as Set Processors When Swix asks if sequence-to-sequence modeling and attention are obsolete for world models, Justin directly corrects the technical premise by explaining that transformers are natively permutation-equivariant set processors rather than sequence models. Swix accepts the clarification as Justin walks through positional embeddings and token-level operations.1:12–3:30 · Guest disagreement 0/10 Origins of World Labs and Founder Backgrounds The hosts open with light background questions asking how Fei-Fei Li and Justin Johnson started World Labs. The dynamic is purely conversational, warm, and collegial with no friction or substantive debate.3:31–13:35 · Guest disagreement 2/10 Compute Scaling, Academic Research, and Future Hardware Alessio and Swix probe whether academic open challenges like ImageNet still work under modern compute and commercial pressures. Fei-Fei clarifies that the issue is academic under-resourcing rather than open versus closed research, while Justin pushes back slightly against the 'hardware lottery' assumption by analyzing GPU performance per watt.13:36–20:41 · Guest disagreement 1/10 The Evolution of Image Captioning and Dense Captioning The guests recount the historical lineage of neural image captioning and dense captioning from their early Stanford lab days. The hosts act primarily as engaged facilitators asking technical follow-ups about forward passes and real-time execution.20:42–29:56 · Guest disagreement 3/10 Pixel Maximalism, Physical Laws, and Machine Understanding Alessio brings up a technical paper on LLMs failing to represent orbital force vectors, sparking a debate on whether latent models can truly learn causal physical laws versus pattern matching. Fei-Fei and Justin draw clear philosophical distinctions between statistical pattern fitting and true human-style understanding.29:56–42:11 · Guest disagreement 2/10 Marble Architecture, Gaussian Splats, and Practical Applications The hosts inquire into the fundamental data representations underlying Marble and Gaussian splats, asking why embodied robotics wasn't emphasized. Fei-Fei clarifies that simulation for robotics is already featured prominently in their plans, explaining the bridge between synthetic data generation and embodied learning.42:12–55:50 · Guest disagreement 3/10 Defining Spatial Intelligence, Embodiment, and Theory Building Fei-Fei immediately rejects the common industry framing of 'a data center full of Einsteins', reframing intelligence as multimodal and highlighting how spatial perception took 540 million years of biological evolution. Alessio counters by noting how language formalizes physical laws like Newton's, leading Justin to separate embodied experience from symbolic theory building.55:50–58:05 · Guest disagreement 4/10 Re-evaluating Model Architectures: Transformers as Set Processors When Swix asks if sequence-to-sequence modeling and attention are obsolete for world models, Justin directly corrects the technical premise by explaining that transformers are natively permutation-equivariant set processors rather than sequence models. Swix accepts the clarification as Justin walks through positional embeddings and token-level operations.1:12–3:30 · The hosts pushing back 0/10 Origins of World Labs and Founder Backgrounds The hosts open with light background questions asking how Fei-Fei Li and Justin Johnson started World Labs. The dynamic is purely conversational, warm, and collegial with no friction or substantive debate.3:31–13:35 · The hosts pushing back 3/10 Compute Scaling, Academic Research, and Future Hardware Alessio and Swix probe whether academic open challenges like ImageNet still work under modern compute and commercial pressures. Fei-Fei clarifies that the issue is academic under-resourcing rather than open versus closed research, while Justin pushes back slightly against the 'hardware lottery' assumption by analyzing GPU performance per watt.13:36–20:41 · The hosts pushing back 1/10 The Evolution of Image Captioning and Dense Captioning The guests recount the historical lineage of neural image captioning and dense captioning from their early Stanford lab days. The hosts act primarily as engaged facilitators asking technical follow-ups about forward passes and real-time execution.20:42–29:56 · The hosts pushing back 3/10 Pixel Maximalism, Physical Laws, and Machine Understanding Alessio brings up a technical paper on LLMs failing to represent orbital force vectors, sparking a debate on whether latent models can truly learn causal physical laws versus pattern matching. Fei-Fei and Justin draw clear philosophical distinctions between statistical pattern fitting and true human-style understanding.29:56–42:11 · The hosts pushing back 2/10 Marble Architecture, Gaussian Splats, and Practical Applications The hosts inquire into the fundamental data representations underlying Marble and Gaussian splats, asking why embodied robotics wasn't emphasized. Fei-Fei clarifies that simulation for robotics is already featured prominently in their plans, explaining the bridge between synthetic data generation and embodied learning.42:12–55:50 · The hosts pushing back 3/10 Defining Spatial Intelligence, Embodiment, and Theory Building Fei-Fei immediately rejects the common industry framing of 'a data center full of Einsteins', reframing intelligence as multimodal and highlighting how spatial perception took 540 million years of biological evolution. Alessio counters by noting how language formalizes physical laws like Newton's, leading Justin to separate embodied experience from symbolic theory building.55:50–58:05 · The hosts pushing back 1/10 Re-evaluating Model Architectures: Transformers as Set Processors When Swix asks if sequence-to-sequence modeling and attention are obsolete for world models, Justin directly corrects the technical premise by explaining that transformers are natively permutation-equivariant set processors rather than sequence models. Swix accepts the clarification as Justin walks through positional embeddings and token-level operations.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 42:36 Fei-Fei bluntly rejects 'data center full of Einsteins' framing

Fei-Fei immediately and flatly dismisses the popular industry soundbite cited by the host, refusing the premise and insisting on a rigorous evolutionary and psychological definition of intelligence.

Hardest push from the hosts ▶ 45:31 Alessio challenges pure spatial primacy using Newtonian formalism

Alessio pushes back against downplaying language by pointing out that language formalizes empirical spatial phenomena into actionable laws like gravity, forcing the guests to address theory building.

Biggest teaching moment ▶ 56:58 Justin educates on transformer architectural foundations as set processors

Justin fundamentally re-educates the host on deep learning theory, correcting the assumption that transformers are sequence models by demonstrating they are permutation-equivariant set models where sequence order is merely injected via embeddings.

The host holds their own ▶ 22:51 Alessio cites Harvard orbital vector paper to challenge LLM world understanding

Alessio demonstrates strong domain mastery by introducing an empirical Harvard study showing LLMs predict planetary orbits without capturing underlying physical force vectors.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Origins of World Labs and Founder Backgrounds 3100 The hosts open with light background questions asking how Fei-Fei Li and Justin Johnson started World Labs. The dynamic is purely conversational, warm, and collegial with no friction or substantive debate.
Compute Scaling, Academic Research, and Future Hardware 5423 Alessio and Swix probe whether academic open challenges like ImageNet still work under modern compute and commercial pressures. Fei-Fei clarifies that the issue is academic under-resourcing rather than open versus closed research, while Justin pushes back slightly against the 'hardware lottery' assumption by analyzing GPU performance per watt.
The Evolution of Image Captioning and Dense Captioning 4311 The guests recount the historical lineage of neural image captioning and dense captioning from their early Stanford lab days. The hosts act primarily as engaged facilitators asking technical follow-ups about forward passes and real-time execution.
Pixel Maximalism, Physical Laws, and Machine Understanding 6533 Alessio brings up a technical paper on LLMs failing to represent orbital force vectors, sparking a debate on whether latent models can truly learn causal physical laws versus pattern matching. Fei-Fei and Justin draw clear philosophical distinctions between statistical pattern fitting and true human-style understanding.
Marble Architecture, Gaussian Splats, and Practical Applications 5422 The hosts inquire into the fundamental data representations underlying Marble and Gaussian splats, asking why embodied robotics wasn't emphasized. Fei-Fei clarifies that simulation for robotics is already featured prominently in their plans, explaining the bridge between synthetic data generation and embodied learning.
Defining Spatial Intelligence, Embodiment, and Theory Building 6633 Fei-Fei immediately rejects the common industry framing of 'a data center full of Einsteins', reframing intelligence as multimodal and highlighting how spatial perception took 540 million years of biological evolution. Alessio counters by noting how language formalizes physical laws like Newton's, leading Justin to separate embodied experience from symbolic theory building.
Re-evaluating Model Architectures: Transformers as Set Processors 4841 When Swix asks if sequence-to-sequence modeling and attention are obsolete for world models, Justin directly corrects the technical premise by explaining that transformers are natively permutation-equivariant set processors rather than sequence models. Swix accepts the clarification as Justin walks through positional embeddings and token-level operations.

Statements from this episode (15)

Disclosure
Li: World Labs was founded to build spatial intelligence beyond LLMs
“I think around, you know, more than two years ago, for sure, I think both independently, both of us have been looking at the development of the large models and thinking about What's beyond language models and this idea of building world models, spatial intell…”
Fei-Fei Li Nov 25, 2025 ▶ 2:14
Assertion Supported
Johnson: AI compute per model has scaled one million-fold since 2012
“And if you think about, you know, AlexNet required this jump from CPUs to GPUs, but even from AlexNet to today, we're getting about a thousand times more performance per card than we had in AlexNet days. And now it's common to train models, not just on one GPU…”
Justin Johnson Nov 25, 2025 ▶ 4:13
Assertion Not checkable as stated
Johnson: Academic labs can no longer train state-of-the-art AI on few GPUs
“Like five or 10 years ago, you really could train state-of-the-art models in the lab even with just a couple of GPUs. But, you know, because that technology was so successful and scaled up so much, then you can't train state-of-the-art models with a couple of …”
Justin Johnson Nov 25, 2025 ▶ 9:51
Assertion Contradicted
Johnson: Nvidia Blackwell offers roughly same performance per watt as Hopper
“Like, if you look at the numbers, like, even going from Hopper to Blackwell, like, the performance per watt is about the same. They mostly make the number of transistors go up, and they make the chip size go up, and they make the power usage go up. But even fr…”
Justin Johnson Nov 25, 2025 ▶ 13:01
Opinion
Li: 3D and 4D spatial structure is fundamentally unlike 1D text
“I do think the architecture of these generative models will share a lot of shareable components, but I think the deeply three D four D spatial world has a level Of structure that is fundamentally different from a purely generative signal that is one dimensiona…”
Fei-Fei Li Nov 25, 2025 ▶ 21:10
Insight
Johnson: Pixels offer a more lossless world representation than tokenized text
“And then like you actually lose something if you translate to this like purely tokenized representations that we use in LLMs, right? Like you lose the font, you lose the line breaks, you lose sort of the two D arrangement on the page. And for a lot of cases, f…”
Justin Johnson Nov 25, 2025 ▶ 22:11
Opinion
Li: Scaling can bridge visual generative AI and physical engineering
“I think this is a matter of scaling data and bettering model. I don't think there's anything fundamental that separates these two.”
Fei-Fei Li Nov 25, 2025 ▶ 27:58
Assertion Not checkable as stated
Li: Marble is the first public high-fidelity 3D generative world model
“It's the first in-class model in the world that generates three D worlds in this level of fidelity that is in the hands of the public.”
Fei-Fei Li Nov 25, 2025 ▶ 31:06
Assertion Supported
Johnson: Gaussian splats render in real time on nearly any client device
“Gaussian splats are really cool because you can render them in real time really efficiently. So you can render on your iPhone, render, render everything. And that's how we get that sort of precise camera control because The splats can be rendered real time on …”
Justin Johnson Nov 25, 2025 ▶ 34:46
Opinion
Li: Marble has real potential to generate synthetic data for robotics
“Marble actually is a, Really potential for helping to generate these synthetic simulated worlds for embodied agent training.”
Fei-Fei Li Nov 25, 2025 ▶ 40:13
Assertion Supported
Li: Evolution spent 540M years on spatial intelligence versus 0.5M on language
“And in nature, you know, it took five hundred, forty million years to optimize perception and spatial intelligence and language in the most Generous estimation of language development is probably half a million years.”
Fei-Fei Li Nov 25, 2025 ▶ 48:47
Insight
Li: LLMs can predict physical trajectories without abstracting physical laws
“I wouldn't be surprised that given enough celestial movement data, an LLM would actually predict pretty accurate movement trajectories... I wouldn't be surprised, but F equals MA or, you know, action equals reaction. That's just a whole different abstraction l…”
Fei-Fei Li Nov 25, 2025 ▶ 52:48
Insight
Johnson: Physical theory-building stems from interactive falsification, not model modality
“Because we're constantly interacting with the world, we're constantly having to build theories about what's happening in the world around us, and then falsify or add evidence to those theories. And I think that that kind of process writ large and scaled up is …”
Justin Johnson Nov 25, 2025 ▶ 54:35
Prediction Not checkable as stated
Li: World models will move beyond sequence-to-sequence architectures
“I think sequence to sequence is actually in world models, I think we are going to see algorithm or architecture beyond sequence to sequence.”
Fei-Fei Li Nov 25, 2025 ▶ 56:27
Insight
Johnson: Transformers are natively models of sets, not sequences
“Transformers are actually not a model of sequences. A transformer is natively a model of sets.”
Justin Johnson Nov 25, 2025 ▶ 56:45
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.