Mostafa Dehghani, research scientist at Google DeepMind, addresses the industry narrative that AI pre-training has hit a wall or reached diminishing returns.
“The way that we used to do pre-training, maybe, like, you know, like two, a year ago or two years ago maybe, like, you know, diminishing return is, like, obvious, but I can see how new ideas are bringing, like, you know, fresh, fresh energy into the pre-training and suddenly just open a door toward, like something exotic that might actually drastically change the base model capability over time.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Mostafa Dehghani
PredictionNot checkable as stated
Fully automated AI self-improvement will eliminate human bottlenecks and trigger breakthroughs
“The moment that we had this full automation, I would say we can close the loop of self-improvement and then it becomes the Like, you know, the problems become like, you know, mostly providing compute for these models to actually do what they want to do. And as…”
Mostafa DehghaniApr 2, 2026▶ 7:07AI is Already Building AI — Google DeepMind’s Mostafa Dehghani
PredictionNot checkable as stated
AI model progress will alternate between pre-training and post-training breakthroughs
“We're going to be having a bit of a swing back and forth between pre-training and post-training.”
Mostafa DehghaniApr 2, 2026▶ 26:55AI is Already Building AI — Google DeepMind’s Mostafa Dehghani
Insight
Post-training techniques cannot compensate for a weak base AI model
“Pre-training is still the foundation and like, you can never post-train your way out of a week-based model.”
Mostafa DehghaniApr 2, 2026▶ 27:01AI is Already Building AI — Google DeepMind’s Mostafa Dehghani
Insight
Video data conveys physical world knowledge to AI more efficiently than text
“So because of that, like picking up a lot of knowledge about the word through language is just not really efficient. I don't want to say that it's impossible, but it's not efficient, you know, like to learn about gravity. If you kind of like, you know, have yo…”
Mostafa DehghaniApr 2, 2026▶ 47:19AI is Already Building AI — Google DeepMind’s Mostafa Dehghani
Insight
Demonstrating that image training lowers text perplexity remains extremely difficult
“So it turned out to be a really, really good model, but it was like really hard to see that. Wow. You know, I train on images and then like Text perplexity goes down. That was hard to see. You know, like the fact that, you know, you train in native model and i…”
Mostafa DehghaniApr 2, 2026▶ 48:54AI is Already Building AI — Google DeepMind’s Mostafa Dehghani
Insight
Jagged intelligence in AI is a structural flaw, not a patchable bug
“Not easy to pinpoint like specific things, but again, like, you know, this is just like my personal opinion and maybe I have colleagues and like the other people like sharing this with me, but I think we're underestimating how hard like jagged intelligence is …”
Mostafa DehghaniApr 2, 2026▶ 54:59AI is Already Building AI — Google DeepMind’s Mostafa Dehghani
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 400 conversations transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.