Manning: Inferring actions from passive observational video is unproven at scale
Chris Manning · Moonlake: Interactive, Multimodal World Models — with Chris Manning and Fan-yun Sun · Apr 2, 2026 · at 9:29
Stanford professor Chris Manning explains why training action-conditioned world models on passive web video is substantially harder than using simulation or action-labeled datasets.
“What's really essential is understanding the consequences of actions, producing an action-conditioned world model, and if you're simply collecting observational video data, which is the easy stuff to collect when you're sort of mining online videos, you don't actually know the actions that are being taken to see how the video is changing, and so if you're never collecting Directly actions, and you're having to try and infer them from what happened in the observed video. That's not impossible, but it's very hard, and it's not really established that you can get that to work at any scale yet”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →