Everything Chris Manning said on any show that made the record, most notable first. Each card names its show and opens the statement there.
Manning: Vision understanding stalled; language does 90% of work in VLMs
“I mean, I think it's fair to say that, you know, vision understanding sort of stalled out, right? You got to object recognition, and then progress just wasn't being made, right? If you look at any of these vision language models, it's the language that's doing…”
Manning: Yann LeCun underestimates language and symbolic representations in intelligence
“Jan LeCun is a dear friend of mine but he has never appreciated the power of language in particular or symbolic representations in general. Yarn is a very visual thinker. He always wants to claim that he thinks visually, and there are no words, symbols, or mat…”
Manning: Transformer internal weights can act as joint representations for world models
“I'm not actually convinced that's right, because although the token production is this autoregressive process that's heading, you know, left to right, I guess don't have to be left or right, but anyway, in sequence of tokens, we could have right to left Arabic…”
Manning: Mainstream vision models fail by operating solely on pixel surfaces
“Believing that there can be a really rich connection between a more symbolic layer of abstracted understanding of visual domains, which aren't in the mainstream vision models, which are still trying to operate on the surface level of pixels.”
Manning: True World Models Require Action Conditioning and Semantic Abstraction
“You only actually have a world model if you can predict, given some action is taken, what is going to change in the world because of that, and in particular that becomes hard over longer time scales, so if you're simply, you know, trying to predict the next vi…”
Manning: Semantic abstractions require five orders of magnitude less data than pixels
“If there are ways in which you can work with five orders of magnitude, less data than people working purely from pixels, you're going to be able to make a lot more progress, a lot more quickly, and that's the bet here.”
Manning: OpenAI's Sora cannot produce compelling gameplay or persistent mechanics
“Don't think you can take Sora and produce compelling gameplay, right? If you want to have a world that you can wander around in a bit, you're good, but what are your abilities to have gameplay mechanics implemented the way you'd like them to be, and to have th…”
Manning: Video Models Lack Genuine 3D Spatial Understanding and Causality
“The reality is that although the visuals do look fantastic, those visuals actually aren't accompanied by an understanding of the three-d world, understanding how objects can move, what the consequences of different actions are, and that's what's really needed …”
Manning: Inferring actions from passive observational video is unproven at scale
“What's really essential is understanding the consequences of actions, producing an action-conditioned world model, and if you're simply collecting observational video data, which is the easy stuff to collect when you're sort of mining online videos, you don't …”
Manning: Reward hacking is unsolved in symbolic and pixel-based models
“I mean, to the extent that there's a misspecified reward that it seems like it could be hacked
In a more symbolic world or in a more pixel based world. I don't know if Sun's got any thoughts, but I don't think that's really being solved.”
Manning: Controlling world models requires both text and visual prompts
“I think it's a mixture. I mean, yeah. I mean, there's clearly a visual component of this and it's not that You know, everything can be text, because of course you want to give a visual look, but there's also a massive amount of giving the overall picture of th…”
Manning: Generative AI video models lack true world-model audio integration
“And whereas in general for the Gen AI video models, there's no actual integration across to audio at all, right? That someone might stick some music or stick a soundscape or whatever else on top of their video so it's not a silent video, but They're in no way …”