Q So from the image to the video, the audio, and then eventually the real world with robotics, And a real world model, because if you can make the image, you, and you can train the model, that means by default, you understand the world. In order to make a video of the world, you have to understand the world, yeah? And the objects in it?
A I think that's, yeah, I think that's like a really good, like, way to think about it. It's like, it's like an intuitive way, uh, to, Interact with the world, right? Like, I would say there's, like, these, like, complementary forms of intelligence, ultimately. There's, like, intuitive intelligence, and then there's, like, a deep reasoning layer. Now, ultimately, you need, for, like, a kind of, like, complete form, you need both, um, and you need them to interact, and I think, like, we've been approaching it more from, like, the intuitive side. Um, images is, like, a very natural way to approach this whole field, because it's not as computationally intensive as, let's say, video, right? But now, yeah, I think, like, we're combining it, it's converging into, like, a multimodal model, and, uh, yeah, we see, like, exactly, like, pre-training on videos gives, like, implicit understanding of the physics of interactions with the real world, and then you can get stuff like action prediction, like robotics out of the same model.
AI assessment note: “pre-training on videos gives, like, implicit understanding of the physics of interactions”