Video data conveys physical world knowledge to AI more efficiently than text
Mostafa Dehghani · AI is Already Building AI — Google DeepMind’s Mostafa Dehghani · Apr 2, 2026 · at 47:19
Mostafa Dehghani, AI research scientist at Google DeepMind, discusses linguistic reporting bias and why visual world models are necessary for AI comprehension.
“So because of that, like picking up a lot of knowledge about the word through language is just not really efficient. I don't want to say that it's impossible, but it's not efficient, you know, like to learn about gravity. If you kind of like, you know, have your model train on videos, it's much easier to get the model to learn about gravity because it just happens in a video than training your model on, on all the textbook to kind of like learn about the concept of gravity, you know, or what is actually gravity.”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →