Assertion certainty 4/5 debate potential 3/5

Manning: Inferring actions from passive observational video is unproven at scale

Chris Manning · Moonlake: Interactive, Multimodal World Models — with Chris Manning and Fan-yun Sun · Apr 2, 2026 · at 9:29

Stanford professor Chris Manning explains why training action-conditioned world models on passive web video is substantially harder than using simulation or action-labeled datasets.

0:00 / 0:41exact quote · 41.7s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“What's really essential is understanding the consequences of actions, producing an action-conditioned world model, and if you're simply collecting observational video data, which is the easy stuff to collect when you're sort of mining online videos, you don't actually know the actions that are being taken to see how the video is changing, and so if you're never collecting Directly actions, and you're having to try and infer them from what happened in the observed video. That's not impossible, but it's very hard, and it's not really established that you can get that to work at any scale yet”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Chris Manning

Opinion
Manning: Vision understanding stalled; language does 90% of work in VLMs
“I mean, I think it's fair to say that, you know, vision understanding sort of stalled out, right? You got to object recognition, and then progress just wasn't being made, right? If you look at any of these vision language models, it's the language that's doing…”
Chris Manning Apr 2, 2026 ▶ 5:02 Moonlake: Interactive, Multimodal World Models — with Chris Manning and Fan-yun Sun
Opinion
Manning: Yann LeCun underestimates language and symbolic representations in intelligence
“Jan LeCun is a dear friend of mine but he has never appreciated the power of language in particular or symbolic representations in general. Yarn is a very visual thinker. He always wants to claim that he thinks visually, and there are no words, symbols, or mat…”
Chris Manning Apr 2, 2026 ▶ 16:20 Moonlake: Interactive, Multimodal World Models — with Chris Manning and Fan-yun Sun
Opinion
Manning: Transformer internal weights can act as joint representations for world models
“I'm not actually convinced that's right, because although the token production is this autoregressive process that's heading, you know, left to right, I guess don't have to be left or right, but anyway, in sequence of tokens, we could have right to left Arabic…”
Chris Manning Apr 2, 2026 ▶ 20:57 Moonlake: Interactive, Multimodal World Models — with Chris Manning and Fan-yun Sun
Insight
Manning: Mainstream vision models fail by operating solely on pixel surfaces
“Believing that there can be a really rich connection between a more symbolic layer of abstracted understanding of visual domains, which aren't in the mainstream vision models, which are still trying to operate on the surface level of pixels.”
Chris Manning Apr 2, 2026 ▶ 5:28 Moonlake: Interactive, Multimodal World Models — with Chris Manning and Fan-yun Sun
Insight
Manning: True World Models Require Action Conditioning and Semantic Abstraction
“You only actually have a world model if you can predict, given some action is taken, what is going to change in the world because of that, and in particular that becomes hard over longer time scales, so if you're simply, you know, trying to predict the next vi…”
Chris Manning Apr 2, 2026 ▶ 7:56 Moonlake: Interactive, Multimodal World Models — with Chris Manning and Fan-yun Sun
Insight
Manning: Semantic abstractions require five orders of magnitude less data than pixels
“If there are ways in which you can work with five orders of magnitude, less data than people working purely from pixels, you're going to be able to make a lot more progress, a lot more quickly, and that's the bet here.”
Chris Manning Apr 2, 2026 ▶ 11:33 Moonlake: Interactive, Multimodal World Models — with Chris Manning and Fan-yun Sun
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.