Sep 4, 2026 · 43m · a16z

Why Fei-Fei Li Is Betting on Spatial Intelligence

Justin Johnson · 13m spoken Ben Mildenhall · 11m spoken Fei-Fei Li · 9m spoken Martin Casado · 5m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of The a16z Show, World Labs co-founders Fei-Fei Li, Justin Johnson, and Ben Mildenhall join host Martin Casado to unveil Atlas, their pioneering world model that unifies 3D reconstruction and generative visual AI through new view prediction. They explore how grounding pixels in spatial geometry revolutionizes creative production, powers real-to-sim robotics pipelines, and serves as an evolutionary prerequisite for artificial general intelligence.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The host as informed peer 3.1 Guest teaching 3.5 Guest disagreement 1.0 The host pushing back 1.5
05100:0015:0030:000:51–5:15 · The host as informed peer 3/10 Introducing Atlas: Core Capabilities and Bullet-Time Demos Casado sets the stage with open, foundational questions about Atlas, asking how it fundamentally differs from standard video models. Johnson and Mildenhall explain new view prediction and spatial context collaboratively without friction.5:15–8:15 · The host as informed peer 2/10 Joint Reconstruction and Generation in a Unified Architecture Casado asks whether Atlas is simply scaled-up video tech or a novel architecture and openly admits he is unfamiliar with 3D representations. Johnson and Li explain depth maps, camera poses, and the historical bifurcation between generation and reconstruction.8:15–10:47 · The host as informed peer 3/10 Defining Spatial Intelligence and Grounded Geometry Casado asks the guests to define spatial intelligence and explain why viewpoint prediction represents a critical milestone. Li details the progression from pixels to geometry, physics, and downstream action planning.10:47–13:15 · The host as informed peer 3/10 The Marble Predecessor and Establishing Spatial Scaling Laws Casado probes why World Labs released Marble before Atlas, prompting Johnson to detail the constraints of Gaussian splats as an output bottleneck. Li notes that prior to their work, no scaling law existed for spatial intelligence.13:15–21:26 · The host as informed peer 3/10 Sparse vs Dense Reconstruction and Visual Context Scaling Casado invites Mildenhall to contrast dense and sparse reconstruction, asking simple questions to clarify why dense capture is so arduous. Mildenhall and Johnson explain the marriage of triangulation with generative hole-filling across context windows.21:26–24:41 · The host as informed peer 3/10 Conviction in Scaling and the Garden Table Breakthrough Casado asks whether the team was truly confident and questions if the architecture is nearing the end of its scaling curve. Li firmly rejects this premise, stating they are only at the beginning, before recounting the breakthrough garden table capture.24:41–29:02 · The host as informed peer 3/10 Creative and Industrial Use Cases: Statefulness and Architecture Casado prompts Mildenhall to explain how Atlas serves creative industries, architecture, and design. Mildenhall walks through traditional 3D design bottlenecks and the necessity of persistent, stateful environments.29:02–35:22 · The host as informed peer 3/10 Robotics, Real-to-Sim Data Generation, and Neural Simulators Casado points out that the connection between a visual world model and robotics is not obvious. Li and Johnson walk him through real-to-sim pipelines, domain randomization, and using learned neural simulators directly as action planners.35:22–40:57 · The host as informed peer 4/10 Temporal Dynamics, 4D Video, and Model Editability Casado relays external expert criticism that the model lacks dynamics and pushes on whether motion conflicts with 3D reconstruction accuracy. Johnson explains that exposure to dynamics during pre-training improves static reconstructions, while Li clarifies control as editability.40:57–43:20 · The host as informed peer 4/10 AI Completeness and the Evolutionary Mandate of Movement Johnson introduces AI completeness, and when Casado offers a colloquial LLM definition, Johnson directly corrects him using Turing completeness and NP reduction to 3-SAT. Casado quickly adapts the premise to visual framing, and Li connects new viewpoints to evolutionary biology.0:51–5:15 · Guest teaching 2/10 Introducing Atlas: Core Capabilities and Bullet-Time Demos Casado sets the stage with open, foundational questions about Atlas, asking how it fundamentally differs from standard video models. Johnson and Mildenhall explain new view prediction and spatial context collaboratively without friction.5:15–8:15 · Guest teaching 4/10 Joint Reconstruction and Generation in a Unified Architecture Casado asks whether Atlas is simply scaled-up video tech or a novel architecture and openly admits he is unfamiliar with 3D representations. Johnson and Li explain depth maps, camera poses, and the historical bifurcation between generation and reconstruction.8:15–10:47 · Guest teaching 3/10 Defining Spatial Intelligence and Grounded Geometry Casado asks the guests to define spatial intelligence and explain why viewpoint prediction represents a critical milestone. Li details the progression from pixels to geometry, physics, and downstream action planning.10:47–13:15 · Guest teaching 3/10 The Marble Predecessor and Establishing Spatial Scaling Laws Casado probes why World Labs released Marble before Atlas, prompting Johnson to detail the constraints of Gaussian splats as an output bottleneck. Li notes that prior to their work, no scaling law existed for spatial intelligence.13:15–21:26 · Guest teaching 4/10 Sparse vs Dense Reconstruction and Visual Context Scaling Casado invites Mildenhall to contrast dense and sparse reconstruction, asking simple questions to clarify why dense capture is so arduous. Mildenhall and Johnson explain the marriage of triangulation with generative hole-filling across context windows.21:26–24:41 · Guest teaching 4/10 Conviction in Scaling and the Garden Table Breakthrough Casado asks whether the team was truly confident and questions if the architecture is nearing the end of its scaling curve. Li firmly rejects this premise, stating they are only at the beginning, before recounting the breakthrough garden table capture.24:41–29:02 · Guest teaching 2/10 Creative and Industrial Use Cases: Statefulness and Architecture Casado prompts Mildenhall to explain how Atlas serves creative industries, architecture, and design. Mildenhall walks through traditional 3D design bottlenecks and the necessity of persistent, stateful environments.29:02–35:22 · Guest teaching 4/10 Robotics, Real-to-Sim Data Generation, and Neural Simulators Casado points out that the connection between a visual world model and robotics is not obvious. Li and Johnson walk him through real-to-sim pipelines, domain randomization, and using learned neural simulators directly as action planners.35:22–40:57 · Guest teaching 4/10 Temporal Dynamics, 4D Video, and Model Editability Casado relays external expert criticism that the model lacks dynamics and pushes on whether motion conflicts with 3D reconstruction accuracy. Johnson explains that exposure to dynamics during pre-training improves static reconstructions, while Li clarifies control as editability.40:57–43:20 · Guest teaching 5/10 AI Completeness and the Evolutionary Mandate of Movement Johnson introduces AI completeness, and when Casado offers a colloquial LLM definition, Johnson directly corrects him using Turing completeness and NP reduction to 3-SAT. Casado quickly adapts the premise to visual framing, and Li connects new viewpoints to evolutionary biology.0:51–5:15 · Guest disagreement 1/10 Introducing Atlas: Core Capabilities and Bullet-Time Demos Casado sets the stage with open, foundational questions about Atlas, asking how it fundamentally differs from standard video models. Johnson and Mildenhall explain new view prediction and spatial context collaboratively without friction.5:15–8:15 · Guest disagreement 1/10 Joint Reconstruction and Generation in a Unified Architecture Casado asks whether Atlas is simply scaled-up video tech or a novel architecture and openly admits he is unfamiliar with 3D representations. Johnson and Li explain depth maps, camera poses, and the historical bifurcation between generation and reconstruction.8:15–10:47 · Guest disagreement 0/10 Defining Spatial Intelligence and Grounded Geometry Casado asks the guests to define spatial intelligence and explain why viewpoint prediction represents a critical milestone. Li details the progression from pixels to geometry, physics, and downstream action planning.10:47–13:15 · Guest disagreement 0/10 The Marble Predecessor and Establishing Spatial Scaling Laws Casado probes why World Labs released Marble before Atlas, prompting Johnson to detail the constraints of Gaussian splats as an output bottleneck. Li notes that prior to their work, no scaling law existed for spatial intelligence.13:15–21:26 · Guest disagreement 0/10 Sparse vs Dense Reconstruction and Visual Context Scaling Casado invites Mildenhall to contrast dense and sparse reconstruction, asking simple questions to clarify why dense capture is so arduous. Mildenhall and Johnson explain the marriage of triangulation with generative hole-filling across context windows.21:26–24:41 · Guest disagreement 2/10 Conviction in Scaling and the Garden Table Breakthrough Casado asks whether the team was truly confident and questions if the architecture is nearing the end of its scaling curve. Li firmly rejects this premise, stating they are only at the beginning, before recounting the breakthrough garden table capture.24:41–29:02 · Guest disagreement 0/10 Creative and Industrial Use Cases: Statefulness and Architecture Casado prompts Mildenhall to explain how Atlas serves creative industries, architecture, and design. Mildenhall walks through traditional 3D design bottlenecks and the necessity of persistent, stateful environments.29:02–35:22 · Guest disagreement 1/10 Robotics, Real-to-Sim Data Generation, and Neural Simulators Casado points out that the connection between a visual world model and robotics is not obvious. Li and Johnson walk him through real-to-sim pipelines, domain randomization, and using learned neural simulators directly as action planners.35:22–40:57 · Guest disagreement 2/10 Temporal Dynamics, 4D Video, and Model Editability Casado relays external expert criticism that the model lacks dynamics and pushes on whether motion conflicts with 3D reconstruction accuracy. Johnson explains that exposure to dynamics during pre-training improves static reconstructions, while Li clarifies control as editability.40:57–43:20 · Guest disagreement 3/10 AI Completeness and the Evolutionary Mandate of Movement Johnson introduces AI completeness, and when Casado offers a colloquial LLM definition, Johnson directly corrects him using Turing completeness and NP reduction to 3-SAT. Casado quickly adapts the premise to visual framing, and Li connects new viewpoints to evolutionary biology.0:51–5:15 · The host pushing back 1/10 Introducing Atlas: Core Capabilities and Bullet-Time Demos Casado sets the stage with open, foundational questions about Atlas, asking how it fundamentally differs from standard video models. Johnson and Mildenhall explain new view prediction and spatial context collaboratively without friction.5:15–8:15 · The host pushing back 2/10 Joint Reconstruction and Generation in a Unified Architecture Casado asks whether Atlas is simply scaled-up video tech or a novel architecture and openly admits he is unfamiliar with 3D representations. Johnson and Li explain depth maps, camera poses, and the historical bifurcation between generation and reconstruction.8:15–10:47 · The host pushing back 1/10 Defining Spatial Intelligence and Grounded Geometry Casado asks the guests to define spatial intelligence and explain why viewpoint prediction represents a critical milestone. Li details the progression from pixels to geometry, physics, and downstream action planning.10:47–13:15 · The host pushing back 1/10 The Marble Predecessor and Establishing Spatial Scaling Laws Casado probes why World Labs released Marble before Atlas, prompting Johnson to detail the constraints of Gaussian splats as an output bottleneck. Li notes that prior to their work, no scaling law existed for spatial intelligence.13:15–21:26 · The host pushing back 1/10 Sparse vs Dense Reconstruction and Visual Context Scaling Casado invites Mildenhall to contrast dense and sparse reconstruction, asking simple questions to clarify why dense capture is so arduous. Mildenhall and Johnson explain the marriage of triangulation with generative hole-filling across context windows.21:26–24:41 · The host pushing back 2/10 Conviction in Scaling and the Garden Table Breakthrough Casado asks whether the team was truly confident and questions if the architecture is nearing the end of its scaling curve. Li firmly rejects this premise, stating they are only at the beginning, before recounting the breakthrough garden table capture.24:41–29:02 · The host pushing back 0/10 Creative and Industrial Use Cases: Statefulness and Architecture Casado prompts Mildenhall to explain how Atlas serves creative industries, architecture, and design. Mildenhall walks through traditional 3D design bottlenecks and the necessity of persistent, stateful environments.29:02–35:22 · The host pushing back 2/10 Robotics, Real-to-Sim Data Generation, and Neural Simulators Casado points out that the connection between a visual world model and robotics is not obvious. Li and Johnson walk him through real-to-sim pipelines, domain randomization, and using learned neural simulators directly as action planners.35:22–40:57 · The host pushing back 4/10 Temporal Dynamics, 4D Video, and Model Editability Casado relays external expert criticism that the model lacks dynamics and pushes on whether motion conflicts with 3D reconstruction accuracy. Johnson explains that exposure to dynamics during pre-training improves static reconstructions, while Li clarifies control as editability.40:57–43:20 · The host pushing back 1/10 AI Completeness and the Evolutionary Mandate of Movement Johnson introduces AI completeness, and when Casado offers a colloquial LLM definition, Johnson directly corrects him using Turing completeness and NP reduction to 3-SAT. Casado quickly adapts the premise to visual framing, and Li connects new viewpoints to evolutionary biology.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 0% · guest 100%0:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%33:00 · the host 0% · guest 100%33:00 · the host 0% · guest 100%36:00 · the host 0% · guest 100%36:00 · the host 0% · guest 100%39:00 · the host 0% · guest 100%39:00 · the host 0% · guest 100%42:00 · the host 0% · guest 100%42:00 · the host 0% · guest 100%
Sharpest disagreement ▶ 41:01 Direct correction on AI completeness

Justin Johnson bluntly interrupts Casado's colloquial LLM explanation with 'No, no' and grounds the definition in formal complexity theory.

Hardest push from the host ▶ 36:45 Challenging dynamics versus reconstruction

Casado relays tough external critique about static models and directly questions whether reconstruction and dynamics are inherently at odds.

Biggest teaching moment ▶ 41:01 Complexity theory tutorial

Johnson corrects Casado's informal conception of AI completeness by explaining its formal linkage to Turing completeness and 3-SAT reduction.

The host holds their own ▶ 42:20 Casado builds on the AI complete thought experiment

Casado immediately demonstrates conceptual grasp of Johnson's premise by extending the visual AI completeness example to proof-solving and whiteboard reveals.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
Introducing Atlas: Core Capabilities and Bullet-Time Demos 3211 Casado sets the stage with open, foundational questions about Atlas, asking how it fundamentally differs from standard video models. Johnson and Mildenhall explain new view prediction and spatial context collaboratively without friction.
Joint Reconstruction and Generation in a Unified Architecture 2412 Casado asks whether Atlas is simply scaled-up video tech or a novel architecture and openly admits he is unfamiliar with 3D representations. Johnson and Li explain depth maps, camera poses, and the historical bifurcation between generation and reconstruction.
Defining Spatial Intelligence and Grounded Geometry 3301 Casado asks the guests to define spatial intelligence and explain why viewpoint prediction represents a critical milestone. Li details the progression from pixels to geometry, physics, and downstream action planning.
The Marble Predecessor and Establishing Spatial Scaling Laws 3301 Casado probes why World Labs released Marble before Atlas, prompting Johnson to detail the constraints of Gaussian splats as an output bottleneck. Li notes that prior to their work, no scaling law existed for spatial intelligence.
Sparse vs Dense Reconstruction and Visual Context Scaling 3401 Casado invites Mildenhall to contrast dense and sparse reconstruction, asking simple questions to clarify why dense capture is so arduous. Mildenhall and Johnson explain the marriage of triangulation with generative hole-filling across context windows.
Conviction in Scaling and the Garden Table Breakthrough 3422 Casado asks whether the team was truly confident and questions if the architecture is nearing the end of its scaling curve. Li firmly rejects this premise, stating they are only at the beginning, before recounting the breakthrough garden table capture.
Creative and Industrial Use Cases: Statefulness and Architecture 3200 Casado prompts Mildenhall to explain how Atlas serves creative industries, architecture, and design. Mildenhall walks through traditional 3D design bottlenecks and the necessity of persistent, stateful environments.
Robotics, Real-to-Sim Data Generation, and Neural Simulators 3412 Casado points out that the connection between a visual world model and robotics is not obvious. Li and Johnson walk him through real-to-sim pipelines, domain randomization, and using learned neural simulators directly as action planners.
Temporal Dynamics, 4D Video, and Model Editability 4424 Casado relays external expert criticism that the model lacks dynamics and pushes on whether motion conflicts with 3D reconstruction accuracy. Johnson explains that exposure to dynamics during pre-training improves static reconstructions, while Li clarifies control as editability.
AI Completeness and the Evolutionary Mandate of Movement 4531 Johnson introduces AI completeness, and when Casado offers a colloquial LLM definition, Johnson directly corrects him using Turing completeness and NP reduction to 3-SAT. Casado quickly adapts the premise to visual framing, and Li connects new viewpoints to evolutionary biology.

Statements from this episode (24)

Assertion Partly supported
Johnson: World Labs' Atlas model generates, reconstructs, and simulates the world
“Atlas is our new next generation world model. It has three basic things. It can generate, reconstruct, and simulate the world.”
Justin Johnson Sep 4, 2026 ▶ 1:06
Assertion Supported
Johnson: Atlas reconstructs 3D spaces from up to 100 frames
“You can input one or multiple up to a hundred frames that are views of the real world, and use those to reconstruct the real world, and that reconstruction can take the case either of a novel, a video flying through the space, or an explicit three-D reconstruc…”
Justin Johnson Sep 4, 2026 ▶ 1:27
Assertion Supported
Johnson: Atlas achieves bullet-time effects using only three iPhones
“But now with Atlas, we can do this with just as few, through three cameras. So like no studio capture, no green screen, no expensive calibration. We can literally stick like three cameras on tri, three iPhones on tripods use these to take sort of a video of so…”
Justin Johnson Sep 4, 2026 ▶ 2:18
Assertion Not yet assessed · timeframe Sep 2026
Johnson: No base model has used new view prediction before Atlas
“And this is a really fundamental primitive that we think is super exciting, a super new primitive For base models that no one's ever done before, right?”
Justin Johnson Sep 4, 2026 ▶ 3:01
Assertion Not yet assessed · timeframe Sep 2026
Johnson: Feeding camera poses natively during pre-training is unprecedented
“It also works on camera poses as a native input to the model, which I don't think anyone's ever done at the pre-training phase before.”
Justin Johnson Sep 4, 2026 ▶ 6:16
Assertion Contradicted
Fei-Fei Li: Atlas is the first unification of pixel generation and reconstruction
“It's the first time we have a unification of pixel generation and pixel reconstruction.”
Fei-Fei Li Sep 4, 2026 ▶ 7:19
Insight
Fei-Fei Li: Spatial intelligence requires generation, reasoning, and interaction
“Spatial intelligence eventually must enable us to both generate what the space is, reason within it, and being able to edit and interact within it.”
Fei-Fei Li Sep 4, 2026 ▶ 8:51
Opinion
Fei-Fei Li: Atlas estimates camera pose to capture critical spatial geometry
“And I do believe Atlas is a significant step forward because now with every single frame, you have a, you can generate and estimate an important piece of information, which is the view, viewpoint, the camera pose, and that is the most critical information one …”
Fei-Fei Li Sep 4, 2026 ▶ 9:31
Opinion
Fei-Fei Li: Atlas grounds generative visual models in spatial context
“So, so in the, on the path to spatial intelligence, generating pixels is definitely a early step, which we have seen with What you call it, gazillions of models. But generating pixels that are truly spatially contextualized and grounded is absolutely another m…”
Fei-Fei Li Sep 4, 2026 ▶ 10:02
Assertion Supported
Li: World Labs founded without a scaling law for spatial intelligence
“There was no scaling law of spatial intelligence.”
Fei-Fei Li Sep 4, 2026 ▶ 12:58
Assertion Partly supported
Fei-Fei Li: Atlas reconstructs Stanford Quad aerial view from ground photos
“Anywhere between three to 25 images, you can reconstruct that entire Stanford quad. But the thing is, we had to show it from an aerial view, but every single input image is Ben standing on the ground, taking a picture from the ground. So everything you see are…”
Fei-Fei Li Sep 4, 2026 ▶ 17:01
Insight
Ben Mildenhall: 3D reconstruction is generative AI with a long context window
“Reconstruction is just like generation with a really long context and you put a lot of stuff in it. Yeah. Right. Like that's the way to actually build this continuum where you kind of bridge between those two things.”
Ben Mildenhall Sep 4, 2026 ▶ 19:17
Assertion Not checkable as stated
Ben Mildenhall: Atlas reduces house 3D reconstruction from 2,000 photos to 30
“I'm taking captures I did with 2000 images of a multi room house and taking it down to like 30, 40 inputs and the fly through looks like basically the same. And this is just like totally inconceivable before. And it's all enabled by building this gracefully sc…”
Ben Mildenhall Sep 4, 2026 ▶ 19:51
Assertion Supported
Ben Mildenhall: World Labs' Atlas model operates autoregressively
“You can do this in sequence because it's an autoregressive model. It's up to you, right, to kind of pick and choose what you add interactively into the context as you generate.”
Ben Mildenhall Sep 4, 2026 ▶ 20:41
Opinion
Johnson: Scaling spatial world models is primarily bottlenecked by training compute
“I think we're basically at the beginning, and we're basically limited by compute at this point, right? Like, data is very important, as Fei-Fei likes to point out, but, like, everything has a bottleneck, and I think the main bottleneck on continuing to scale t…”
Justin Johnson Sep 4, 2026 ▶ 22:56
Assertion Not checkable as stated
Johnson: World Labs' models improved significantly with each increase in compute
“Each time we made the model bigger and each time we trained it for longer, each time we put it on more chips, like it got significantly better.”
Justin Johnson Sep 4, 2026 ▶ 23:17
Opinion
Mildenhall: No creative uses a single monolithic AI model for entire tasks
“I don't think there's a single person out there using one monolithic model, not even C dance or whatever for their entire task. People will have this, like, you know, kind of bunch of storyboards and mood boards of images they pull out from, like, their favori…”
Ben Mildenhall Sep 4, 2026 ▶ 25:58
Insight
Mildenhall: Spatial creative workflows require stateful assets over ephemeral prompting
“Like, this statefulness and persistence is so key in how people think about spatial reasoning and, like, developing an environment over time. Like, people don't think in this ephemeral, like, generate a thing, like, just throw it away, keep my text prompts. Li…”
Ben Mildenhall Sep 4, 2026 ▶ 27:01
Opinion
Fei-Fei Li: The biggest problem in robotics is data, not chips yet
“We should zoom out and recognize the biggest problem right now in robotics is actually data. One day it'll be chips, but for now it's data.”
Fei-Fei Li Sep 4, 2026 ▶ 30:50
Assertion Not checkable as stated
Fei-Fei Li: No existing frontier foundation model is robust enough for robotics
“That's just the first part of this, is meeting the robotics needs in the current technology, because we don't yet have a frontier foundation model that's robust enough for robotics.”
Fei-Fei Li Sep 4, 2026 ▶ 31:57
Insight
Li: Spatial intelligence requires closing the loop between seeing and interacting
“Intelligence is not sitting there stuck and just seeing something or interpreting something when it comes to space and physical space, right? It's really this closing the loop between seeing and experiencing an interaction.”
Fei-Fei Li Sep 4, 2026 ▶ 40:39
Opinion
Johnson: Generative new view prediction in Atlas is AI-complete
“But I think that's something we're kind of realizing, and Ben was talking about this earlier today, is, like, new view prediction, this primitive that we have in Atlas, especially generative new view prediction, this is also AI complete.”
Justin Johnson Sep 4, 2026 ▶ 42:09
Insight
Li: Evolution gave animals eyes because movement requires new viewpoint prediction
“That new viewpoint prediction is exactly evolution had to solve by making animals move. You, nature give animals eyes. But nature didn't give trees eyes. Why? Because when you move, you see a new viewpoint.”
Fei-Fei Li Sep 4, 2026 ▶ 42:46
Opinion
Li: Next viewpoint prediction is equivalent to next-token prediction
“So, so we do believe very strongly that next viewpoint prediction is, is the equivalent of next token prediction.”
Fei-Fei Li Sep 4, 2026 ▶ 43:12
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.