Apr 2, 2026 · 1h 4m · mad

AI is Already Building AI — Google DeepMind’s Mostafa Dehghani

Mostafa Dehghani · 52m spoken Matt Turck · 8m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of The MAD Podcast, Google DeepMind research scientist Mostafa Dehghani discusses the shift toward self-improving AI systems, examining test-time compute, autonomous research loops, architectural breakthroughs like Vision Transformers and Nano Banana, and the key bottlenecks remaining in AI evaluation and real-world grounding.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 14.1% of the talking time here. How this is scored →

Matt as informed peer 4.1 Guest teaching 3.5 Guest disagreement 0.3 Matt pushing back 0.8
05100:0015:0030:0045:001:00:001:17–5:04 · Matt as informed peer 3/10 Decoupling AI Looping: Inference Compute versus Development Self-Improvement Matt opens by asking Mostafa to clarify the concept of recursive AI loops. Mostafa provides a structured breakdown separating inference compute looping from development self-improvement.5:04–10:01 · Matt as informed peer 6/10 Recursive Self-Improvement and Autonomous AI Research via AutoResearch Matt cites Karpathy's auto research project and ICLR papers, synthesizing the key distinction between AI assistance tools and direct weight updates. Mostafa validates this framing while noting missing long-horizon capabilities.10:01–12:36 · Matt as informed peer 4/10 Evaluation Bottlenecks and Safe Infrastructure for Self-Improving Agents Matt identifies evaluation as a core roadblock in self-improving systems. Mostafa agrees and explains the philosophical and infrastructure hurdles of running safe evaluation environments.12:36–15:33 · Matt as informed peer 6/10 Formal Verification and Grounding Signals to Prevent Model Collapse Matt references formal verification from a prior guest conversation with Karina Hong of Axiom Math and asks Mostafa to define model collapse. Mostafa notes formal verification works well for code/math but struggles in messier domains.15:33–20:57 · Matt as informed peer 5/10 Navigating the Trade-Off Between Model Specialization and Generalization Matt asks about the tension between model specialization and generalization in self-improving loops. Mostafa explains how post-training overfitting and organizational strategy drive short-term specialization.20:57–23:56 · Matt as informed peer 3/10 Automating AI Researchers and Essential Future Human Skills Matt asks whether top researchers like Karpathy are automating themselves out of a job. Mostafa reframes the question around high-level strategic decision-making rather than pure technical execution.23:56–26:22 · Matt as informed peer 4/10 Redefining Data: Physical Grounding and Multi-Sensory Environment Interaction Matt questions whether data remains essential if AI creates itself, noting the rise of sensors-as-a-service. Mostafa expands the definition of data to include multi-sensory physical grounding.26:26–29:46 · Matt as informed peer 5/10 Balancing Pre-Training and Post-Training Gains in AI Matt directly pushes back against the prevailing narrative that pre-training is dead. Mostafa agrees that pre-training continues to unlock major gains through exotic new recipes.29:46–33:43 · Matt as informed peer 4/10 Defining Continual Learning vs. Self-Improvement Loops Matt prompts Mostafa to define continual learning versus self-improvement loops. Mostafa explains that continual learning focuses on staying current to prevent catastrophic forgetting.33:43–39:57 · Matt as informed peer 5/10 Mostafa Dehghani's Career Journey and the Universal Transformer Matt demonstrates research background knowledge by citing Mostafa's 2018/2019 Universal Transformer paper. Mostafa reflects on working with Lukasz Kaiser and the origins of test-time compute.39:57–43:47 · Matt as informed peer 6/10 The Vision Transformer (ViT) and Unifying Multimodal Architectures Matt recites the exact title of the Vision Transformer paper and synthesizes its role in unifying multimodal architectures. Mostafa details the simple patchify intuition behind ViT.43:47–52:44 · Matt as informed peer 5/10 Nano Banana: Multimodal World Models and Interleaved Generation Matt details Google's Nano Banana release timeline and asks how native multimodality differs from standard text-to-image translation. Mostafa explains reporting bias and interleaved generation planning.52:44–54:53 · Matt as informed peer 4/10 Technical Efficiency and Serving Optimizations in Nano Banana 2 Matt asks about the efficiency gains and architecture behind Nano Banana 2 Flash. Mostafa attributes speedups to model sizing, distillation recipes, and serving infrastructure optimizations.54:58–58:20 · Matt as informed peer 3/10 The Problem of Jagged Intelligence in AI Systems Matt prompts Mostafa for a hot take on what the AI field is getting wrong. Mostafa highlights jagged intelligence as a fundamental structural issue rather than a patchable bug.58:20–1:03:58 · Matt as informed peer 3/10 Overconfidence in Purely Technical Solutions for AI Mostafa rejects the assumption that technical progress alone is sufficient, warning against overconfidence and illustrating agent compounding error rates with concrete math.1:03:58–1:04:10 · Matt as informed peer 0/10 Concluding Remarks and Final Thanks Brief closing pleasantries, sign-off, and audience housekeeping.1:17–5:04 · Guest teaching 4/10 Decoupling AI Looping: Inference Compute versus Development Self-Improvement Matt opens by asking Mostafa to clarify the concept of recursive AI loops. Mostafa provides a structured breakdown separating inference compute looping from development self-improvement.5:04–10:01 · Guest teaching 3/10 Recursive Self-Improvement and Autonomous AI Research via AutoResearch Matt cites Karpathy's auto research project and ICLR papers, synthesizing the key distinction between AI assistance tools and direct weight updates. Mostafa validates this framing while noting missing long-horizon capabilities.10:01–12:36 · Guest teaching 4/10 Evaluation Bottlenecks and Safe Infrastructure for Self-Improving Agents Matt identifies evaluation as a core roadblock in self-improving systems. Mostafa agrees and explains the philosophical and infrastructure hurdles of running safe evaluation environments.12:36–15:33 · Guest teaching 4/10 Formal Verification and Grounding Signals to Prevent Model Collapse Matt references formal verification from a prior guest conversation with Karina Hong of Axiom Math and asks Mostafa to define model collapse. Mostafa notes formal verification works well for code/math but struggles in messier domains.15:33–20:57 · Guest teaching 5/10 Navigating the Trade-Off Between Model Specialization and Generalization Matt asks about the tension between model specialization and generalization in self-improving loops. Mostafa explains how post-training overfitting and organizational strategy drive short-term specialization.20:57–23:56 · Guest teaching 3/10 Automating AI Researchers and Essential Future Human Skills Matt asks whether top researchers like Karpathy are automating themselves out of a job. Mostafa reframes the question around high-level strategic decision-making rather than pure technical execution.23:56–26:22 · Guest teaching 4/10 Redefining Data: Physical Grounding and Multi-Sensory Environment Interaction Matt questions whether data remains essential if AI creates itself, noting the rise of sensors-as-a-service. Mostafa expands the definition of data to include multi-sensory physical grounding.26:26–29:46 · Guest teaching 3/10 Balancing Pre-Training and Post-Training Gains in AI Matt directly pushes back against the prevailing narrative that pre-training is dead. Mostafa agrees that pre-training continues to unlock major gains through exotic new recipes.29:46–33:43 · Guest teaching 4/10 Defining Continual Learning vs. Self-Improvement Loops Matt prompts Mostafa to define continual learning versus self-improvement loops. Mostafa explains that continual learning focuses on staying current to prevent catastrophic forgetting.33:43–39:57 · Guest teaching 3/10 Mostafa Dehghani's Career Journey and the Universal Transformer Matt demonstrates research background knowledge by citing Mostafa's 2018/2019 Universal Transformer paper. Mostafa reflects on working with Lukasz Kaiser and the origins of test-time compute.39:57–43:47 · Guest teaching 3/10 The Vision Transformer (ViT) and Unifying Multimodal Architectures Matt recites the exact title of the Vision Transformer paper and synthesizes its role in unifying multimodal architectures. Mostafa details the simple patchify intuition behind ViT.43:47–52:44 · Guest teaching 4/10 Nano Banana: Multimodal World Models and Interleaved Generation Matt details Google's Nano Banana release timeline and asks how native multimodality differs from standard text-to-image translation. Mostafa explains reporting bias and interleaved generation planning.52:44–54:53 · Guest teaching 3/10 Technical Efficiency and Serving Optimizations in Nano Banana 2 Matt asks about the efficiency gains and architecture behind Nano Banana 2 Flash. Mostafa attributes speedups to model sizing, distillation recipes, and serving infrastructure optimizations.54:58–58:20 · Guest teaching 4/10 The Problem of Jagged Intelligence in AI Systems Matt prompts Mostafa for a hot take on what the AI field is getting wrong. Mostafa highlights jagged intelligence as a fundamental structural issue rather than a patchable bug.58:20–1:03:58 · Guest teaching 5/10 Overconfidence in Purely Technical Solutions for AI Mostafa rejects the assumption that technical progress alone is sufficient, warning against overconfidence and illustrating agent compounding error rates with concrete math.1:03:58–1:04:10 · Guest teaching 0/10 Concluding Remarks and Final Thanks Brief closing pleasantries, sign-off, and audience housekeeping.1:17–5:04 · Guest disagreement 0/10 Decoupling AI Looping: Inference Compute versus Development Self-Improvement Matt opens by asking Mostafa to clarify the concept of recursive AI loops. Mostafa provides a structured breakdown separating inference compute looping from development self-improvement.5:04–10:01 · Guest disagreement 0/10 Recursive Self-Improvement and Autonomous AI Research via AutoResearch Matt cites Karpathy's auto research project and ICLR papers, synthesizing the key distinction between AI assistance tools and direct weight updates. Mostafa validates this framing while noting missing long-horizon capabilities.10:01–12:36 · Guest disagreement 0/10 Evaluation Bottlenecks and Safe Infrastructure for Self-Improving Agents Matt identifies evaluation as a core roadblock in self-improving systems. Mostafa agrees and explains the philosophical and infrastructure hurdles of running safe evaluation environments.12:36–15:33 · Guest disagreement 1/10 Formal Verification and Grounding Signals to Prevent Model Collapse Matt references formal verification from a prior guest conversation with Karina Hong of Axiom Math and asks Mostafa to define model collapse. Mostafa notes formal verification works well for code/math but struggles in messier domains.15:33–20:57 · Guest disagreement 0/10 Navigating the Trade-Off Between Model Specialization and Generalization Matt asks about the tension between model specialization and generalization in self-improving loops. Mostafa explains how post-training overfitting and organizational strategy drive short-term specialization.20:57–23:56 · Guest disagreement 1/10 Automating AI Researchers and Essential Future Human Skills Matt asks whether top researchers like Karpathy are automating themselves out of a job. Mostafa reframes the question around high-level strategic decision-making rather than pure technical execution.23:56–26:22 · Guest disagreement 0/10 Redefining Data: Physical Grounding and Multi-Sensory Environment Interaction Matt questions whether data remains essential if AI creates itself, noting the rise of sensors-as-a-service. Mostafa expands the definition of data to include multi-sensory physical grounding.26:26–29:46 · Guest disagreement 1/10 Balancing Pre-Training and Post-Training Gains in AI Matt directly pushes back against the prevailing narrative that pre-training is dead. Mostafa agrees that pre-training continues to unlock major gains through exotic new recipes.29:46–33:43 · Guest disagreement 0/10 Defining Continual Learning vs. Self-Improvement Loops Matt prompts Mostafa to define continual learning versus self-improvement loops. Mostafa explains that continual learning focuses on staying current to prevent catastrophic forgetting.33:43–39:57 · Guest disagreement 0/10 Mostafa Dehghani's Career Journey and the Universal Transformer Matt demonstrates research background knowledge by citing Mostafa's 2018/2019 Universal Transformer paper. Mostafa reflects on working with Lukasz Kaiser and the origins of test-time compute.39:57–43:47 · Guest disagreement 0/10 The Vision Transformer (ViT) and Unifying Multimodal Architectures Matt recites the exact title of the Vision Transformer paper and synthesizes its role in unifying multimodal architectures. Mostafa details the simple patchify intuition behind ViT.43:47–52:44 · Guest disagreement 0/10 Nano Banana: Multimodal World Models and Interleaved Generation Matt details Google's Nano Banana release timeline and asks how native multimodality differs from standard text-to-image translation. Mostafa explains reporting bias and interleaved generation planning.52:44–54:53 · Guest disagreement 0/10 Technical Efficiency and Serving Optimizations in Nano Banana 2 Matt asks about the efficiency gains and architecture behind Nano Banana 2 Flash. Mostafa attributes speedups to model sizing, distillation recipes, and serving infrastructure optimizations.54:58–58:20 · Guest disagreement 1/10 The Problem of Jagged Intelligence in AI Systems Matt prompts Mostafa for a hot take on what the AI field is getting wrong. Mostafa highlights jagged intelligence as a fundamental structural issue rather than a patchable bug.58:20–1:03:58 · Guest disagreement 1/10 Overconfidence in Purely Technical Solutions for AI Mostafa rejects the assumption that technical progress alone is sufficient, warning against overconfidence and illustrating agent compounding error rates with concrete math.1:03:58–1:04:10 · Guest disagreement 0/10 Concluding Remarks and Final Thanks Brief closing pleasantries, sign-off, and audience housekeeping.1:17–5:04 · Matt pushing back 0/10 Decoupling AI Looping: Inference Compute versus Development Self-Improvement Matt opens by asking Mostafa to clarify the concept of recursive AI loops. Mostafa provides a structured breakdown separating inference compute looping from development self-improvement.5:04–10:01 · Matt pushing back 2/10 Recursive Self-Improvement and Autonomous AI Research via AutoResearch Matt cites Karpathy's auto research project and ICLR papers, synthesizing the key distinction between AI assistance tools and direct weight updates. Mostafa validates this framing while noting missing long-horizon capabilities.10:01–12:36 · Matt pushing back 1/10 Evaluation Bottlenecks and Safe Infrastructure for Self-Improving Agents Matt identifies evaluation as a core roadblock in self-improving systems. Mostafa agrees and explains the philosophical and infrastructure hurdles of running safe evaluation environments.12:36–15:33 · Matt pushing back 2/10 Formal Verification and Grounding Signals to Prevent Model Collapse Matt references formal verification from a prior guest conversation with Karina Hong of Axiom Math and asks Mostafa to define model collapse. Mostafa notes formal verification works well for code/math but struggles in messier domains.15:33–20:57 · Matt pushing back 1/10 Navigating the Trade-Off Between Model Specialization and Generalization Matt asks about the tension between model specialization and generalization in self-improving loops. Mostafa explains how post-training overfitting and organizational strategy drive short-term specialization.20:57–23:56 · Matt pushing back 1/10 Automating AI Researchers and Essential Future Human Skills Matt asks whether top researchers like Karpathy are automating themselves out of a job. Mostafa reframes the question around high-level strategic decision-making rather than pure technical execution.23:56–26:22 · Matt pushing back 1/10 Redefining Data: Physical Grounding and Multi-Sensory Environment Interaction Matt questions whether data remains essential if AI creates itself, noting the rise of sensors-as-a-service. Mostafa expands the definition of data to include multi-sensory physical grounding.26:26–29:46 · Matt pushing back 4/10 Balancing Pre-Training and Post-Training Gains in AI Matt directly pushes back against the prevailing narrative that pre-training is dead. Mostafa agrees that pre-training continues to unlock major gains through exotic new recipes.29:46–33:43 · Matt pushing back 1/10 Defining Continual Learning vs. Self-Improvement Loops Matt prompts Mostafa to define continual learning versus self-improvement loops. Mostafa explains that continual learning focuses on staying current to prevent catastrophic forgetting.33:43–39:57 · Matt pushing back 0/10 Mostafa Dehghani's Career Journey and the Universal Transformer Matt demonstrates research background knowledge by citing Mostafa's 2018/2019 Universal Transformer paper. Mostafa reflects on working with Lukasz Kaiser and the origins of test-time compute.39:57–43:47 · Matt pushing back 0/10 The Vision Transformer (ViT) and Unifying Multimodal Architectures Matt recites the exact title of the Vision Transformer paper and synthesizes its role in unifying multimodal architectures. Mostafa details the simple patchify intuition behind ViT.43:47–52:44 · Matt pushing back 0/10 Nano Banana: Multimodal World Models and Interleaved Generation Matt details Google's Nano Banana release timeline and asks how native multimodality differs from standard text-to-image translation. Mostafa explains reporting bias and interleaved generation planning.52:44–54:53 · Matt pushing back 0/10 Technical Efficiency and Serving Optimizations in Nano Banana 2 Matt asks about the efficiency gains and architecture behind Nano Banana 2 Flash. Mostafa attributes speedups to model sizing, distillation recipes, and serving infrastructure optimizations.54:58–58:20 · Matt pushing back 0/10 The Problem of Jagged Intelligence in AI Systems Matt prompts Mostafa for a hot take on what the AI field is getting wrong. Mostafa highlights jagged intelligence as a fundamental structural issue rather than a patchable bug.58:20–1:03:58 · Matt pushing back 0/10 Overconfidence in Purely Technical Solutions for AI Mostafa rejects the assumption that technical progress alone is sufficient, warning against overconfidence and illustrating agent compounding error rates with concrete math.1:03:58–1:04:10 · Matt pushing back 0/10 Concluding Remarks and Final Thanks Brief closing pleasantries, sign-off, and audience housekeeping.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 36% · guest 64%0:00 · Matt 36% · guest 64%3:00 · Matt 16.7% · guest 83.3%3:00 · Matt 16.7% · guest 83.3%6:00 · Matt 9.1% · guest 90.9%6:00 · Matt 9.1% · guest 90.9%9:00 · Matt 28.8% · guest 71.2%9:00 · Matt 28.8% · guest 71.2%12:00 · Matt 16% · guest 84%12:00 · Matt 16% · guest 84%15:00 · Matt 10.2% · guest 89.8%15:00 · Matt 10.2% · guest 89.8%18:00 · Matt 7.1% · guest 92.9%18:00 · Matt 7.1% · guest 92.9%21:00 · Matt 16.6% · guest 83.4%21:00 · Matt 16.6% · guest 83.4%24:00 · Matt 24.4% · guest 75.6%24:00 · Matt 24.4% · guest 75.6%27:00 · Matt 14.8% · guest 85.2%27:00 · Matt 14.8% · guest 85.2%30:00 · Matt 8.1% · guest 91.9%30:00 · Matt 8.1% · guest 91.9%33:00 · Matt 7% · guest 93%33:00 · Matt 7% · guest 93%36:00 · Matt 6.9% · guest 93.1%36:00 · Matt 6.9% · guest 93.1%39:00 · Matt 9.4% · guest 90.6%39:00 · Matt 9.4% · guest 90.6%42:00 · Matt 45.1% · guest 54.9%42:00 · Matt 45.1% · guest 54.9%45:00 · Matt 2.1% · guest 97.9%45:00 · Matt 2.1% · guest 97.9%48:00 · Matt 0% · guest 100%48:00 · Matt 0% · guest 100%51:00 · Matt 11.3% · guest 88.7%51:00 · Matt 11.3% · guest 88.7%54:00 · Matt 9.2% · guest 90.8%54:00 · Matt 9.2% · guest 90.8%57:00 · Matt 5.3% · guest 94.7%57:00 · Matt 5.3% · guest 94.7%1:00:00 · Matt 4.9% · guest 95.1%1:00:00 · Matt 4.9% · guest 95.1%1:03:00 · Matt 30.5% · guest 69.5%1:03:00 · Matt 30.5% · guest 69.5%
Sharpest disagreement ▶ 58:24 Rejecting technical-only AI progress

Mostafa explicitly disagrees with the widespread optimism that purely technical advancements will automatically solve broader societal, governance, and institutional challenges.

Hardest push from Matt ▶ 28:14 Challenging 'pre-training is dead' narrative

Matt directly confronts the guest with the recent industry consensus that pre-training is dead, pushing Mostafa to explicitly defend his contrasting stance.

Biggest teaching moment ▶ 1:00:45 Compounding probability in multi-step agents

Mostafa educates the host on agent reliability using concrete probability math, demonstrating that 95% accuracy over 100 steps results in less than 1% overall task success.

Matt holds his own ▶ 39:57 Citing exact Vision Transformer paper title

Matt demonstrates deep familiarity with Mostafa's research history by reciting the exact title and patch dimensions of the seminal Vision Transformer paper.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Decoupling AI Looping: Inference Compute versus Development Self-Improvement 3400 Matt opens by asking Mostafa to clarify the concept of recursive AI loops. Mostafa provides a structured breakdown separating inference compute looping from development self-improvement.
Recursive Self-Improvement and Autonomous AI Research via AutoResearch 6302 Matt cites Karpathy's auto research project and ICLR papers, synthesizing the key distinction between AI assistance tools and direct weight updates. Mostafa validates this framing while noting missing long-horizon capabilities.
Evaluation Bottlenecks and Safe Infrastructure for Self-Improving Agents 4401 Matt identifies evaluation as a core roadblock in self-improving systems. Mostafa agrees and explains the philosophical and infrastructure hurdles of running safe evaluation environments.
Formal Verification and Grounding Signals to Prevent Model Collapse 6412 Matt references formal verification from a prior guest conversation with Karina Hong of Axiom Math and asks Mostafa to define model collapse. Mostafa notes formal verification works well for code/math but struggles in messier domains.
Navigating the Trade-Off Between Model Specialization and Generalization 5501 Matt asks about the tension between model specialization and generalization in self-improving loops. Mostafa explains how post-training overfitting and organizational strategy drive short-term specialization.
Automating AI Researchers and Essential Future Human Skills 3311 Matt asks whether top researchers like Karpathy are automating themselves out of a job. Mostafa reframes the question around high-level strategic decision-making rather than pure technical execution.
Redefining Data: Physical Grounding and Multi-Sensory Environment Interaction 4401 Matt questions whether data remains essential if AI creates itself, noting the rise of sensors-as-a-service. Mostafa expands the definition of data to include multi-sensory physical grounding.
Balancing Pre-Training and Post-Training Gains in AI 5314 Matt directly pushes back against the prevailing narrative that pre-training is dead. Mostafa agrees that pre-training continues to unlock major gains through exotic new recipes.
Defining Continual Learning vs. Self-Improvement Loops 4401 Matt prompts Mostafa to define continual learning versus self-improvement loops. Mostafa explains that continual learning focuses on staying current to prevent catastrophic forgetting.
Mostafa Dehghani's Career Journey and the Universal Transformer 5300 Matt demonstrates research background knowledge by citing Mostafa's 2018/2019 Universal Transformer paper. Mostafa reflects on working with Lukasz Kaiser and the origins of test-time compute.
The Vision Transformer (ViT) and Unifying Multimodal Architectures 6300 Matt recites the exact title of the Vision Transformer paper and synthesizes its role in unifying multimodal architectures. Mostafa details the simple patchify intuition behind ViT.
Nano Banana: Multimodal World Models and Interleaved Generation 5400 Matt details Google's Nano Banana release timeline and asks how native multimodality differs from standard text-to-image translation. Mostafa explains reporting bias and interleaved generation planning.
Technical Efficiency and Serving Optimizations in Nano Banana 2 4300 Matt asks about the efficiency gains and architecture behind Nano Banana 2 Flash. Mostafa attributes speedups to model sizing, distillation recipes, and serving infrastructure optimizations.
The Problem of Jagged Intelligence in AI Systems 3410 Matt prompts Mostafa for a hot take on what the AI field is getting wrong. Mostafa highlights jagged intelligence as a fundamental structural issue rather than a patchable bug.
Overconfidence in Purely Technical Solutions for AI 3510 Mostafa rejects the assumption that technical progress alone is sufficient, warning against overconfidence and illustrating agent compounding error rates with concrete math.
Concluding Remarks and Final Thanks 0000 Brief closing pleasantries, sign-off, and audience housekeeping.

Statements from this episode (27)

Assertion Not checkable as stated
AI looping is a top R&D investment for almost every major lab
“Definitely one of the top is active areas for almost every lab to invest in looping”
Mostafa Dehghani Apr 2, 2026 ▶ 1:34
Assertion Not checkable as stated
Every major AI lab uses previous model generations to build new ones
“In almost every lab the new generation of the models are built heavily using the previous generation of the models. I think that's basically like the case again, everywhere.”
Mostafa Dehghani Apr 2, 2026 ▶ 6:05
Prediction Not checkable as stated
Fully automated AI self-improvement will eliminate human bottlenecks and trigger breakthroughs
“The moment that we had this full automation, I would say we can close the loop of self-improvement and then it becomes the Like, you know, the problems become like, you know, mostly providing compute for these models to actually do what they want to do. And as…”
Mostafa Dehghani Apr 2, 2026 ▶ 7:07
Assertion Not checkable as stated
Karpathy's AutoResearch is an early example of AI doing sensible research
“That is definitely. And I think that was one of the early examples of like seeing these models actually doing something super sensible on the research side.”
Mostafa Dehghani Apr 2, 2026 ▶ 7:46
Assertion Not checkable as stated
AI research lacks evaluations to measure progress toward true self-improvement loops
“And the fact that we don't have evals that like, or like even defining evals that, that can maybe measure, oh, not how close we are to the point that we can actually get, get a self-improvement loop. It's just like, we don't have that. And it's just making it …”
Mostafa Dehghani Apr 2, 2026 ▶ 10:49
Disclosure
Safe execution environments are currently the bottleneck for autonomous AI development
“Like in a safe setup, you know, where they can like, because right now we definitely, we don't, we're not confident about, you know, them doing the right things all the time and measuring like how much they can push and how long they can push a task is very di…”
Mostafa Dehghani Apr 2, 2026 ▶ 12:07
Opinion
Formal verification is crucial for AI self-improvement, but not a silver bullet
“In my opinion, formal verification is one of the most powerful like keys to enable like self-improvement, but it's not beaky.”
Mostafa Dehghani Apr 2, 2026 ▶ 12:52
Insight
Anchoring self-improvement loops with real external signals prevents AI model collapse
“Model collapse mainly happens when you have a loop that is Completely closed. Right. And if you don't have any outside signal and just the model, for example, talking to itself or operating in a very like a restricted environment there's a good chance that you…”
Mostafa Dehghani Apr 2, 2026 ▶ 14:14
Insight
Building specialist models is currently the fastest path toward generalist AI
“Short term I would say like building a specialist model is like probably the fastest way to learn like what is actually possible. And in, in many cases, these like specialized model are becoming Stepping stone toward a generalist model, which is like super val…”
Mostafa Dehghani Apr 2, 2026 ▶ 16:51
Prediction Not checkable as stated
Becoming a deep expert in a narrow subject will soon lose value
“Becoming absolute expert about a very specific subject, most likely Is not going to be useful in, in, in like the near future.”
Mostafa Dehghani Apr 2, 2026 ▶ 22:38
Prediction Not checkable as stated
AI data work will shift toward physical world grounding and interactive environments
“At the end of the day, I think, like, the work that we're doing on the data side most likely is going to shift toward building environments or making sure that these models can interact with physical words, and then it becomes more of a problem of, okay, how c…”
Mostafa Dehghani Apr 2, 2026 ▶ 24:49
Prediction Not checkable as stated
AI model progress will alternate between pre-training and post-training breakthroughs
“We're going to be having a bit of a swing back and forth between pre-training and post-training.”
Mostafa Dehghani Apr 2, 2026 ▶ 26:55
Insight
Post-training techniques cannot compensate for a weak base AI model
“Pre-training is still the foundation and like, you can never post-train your way out of a week-based model.”
Mostafa Dehghani Apr 2, 2026 ▶ 27:01
Prediction Not checkable as stated
New pre-training techniques will drastically boost base AI model capabilities
“The way that we used to do pre-training, maybe, like, you know, like two, a year ago or two years ago maybe, like, you know, diminishing return is, like, obvious, but I can see how new ideas are bringing, like, you know, fresh, fresh energy into the pre-traini…”
Mostafa Dehghani Apr 2, 2026 ▶ 29:18
Insight
Self-improvement increases AI capability; continual learning keeps its knowledge fresh
“So, self-improvement is about a model getting smarter. Over time and improving its capability, like the model itself doing it. Continual learning is mostly about a model staying current, right?”
Mostafa Dehghani Apr 2, 2026 ▶ 30:10
Assertion Not checkable as stated
AI continual learning research still lacks a proven recipe for productionization
“Like one side is, I think like the research is not like yet to a very, like to a point that you think that, oh, you know, this is the recipe, you know, I just need to kind of like, you know, exploit it and push productionization.”
Mostafa Dehghani Apr 2, 2026 ▶ 32:04
Disclosure
Dehghani initially thought Google's Transformer architecture was random and would die
“And I was like, I don't know if I want to go with this team. It's just like, they're doing something random. Like who, like everybody's doing LST. I'm like, why should I go and work with like a group of people who are working on this like random architecture, …”
Mostafa Dehghani Apr 2, 2026 ▶ 35:23
Insight
Looping provides parameter-free FLOPS, whereas MoE architectures provide FLOPS-free parameters
“In mixture of experts, you have flops free parameters. So, so parameters that they're not actually bringing any flops. And in, in like looping, you have parameter free flops where you don't have extra parameters for the extra flops that you're throwing on this…”
Mostafa Dehghani Apr 2, 2026 ▶ 39:22
Insight
Vision Transformers benefited heavily from software infrastructure built for language models
“There was also benefit of doing that simply because the rest of the, that the machine learning field, which was working on, on, on language, they were using this, like architecture. So they were building infra for it, making it faster. And, you know, like the,…”
Mostafa Dehghani Apr 2, 2026 ▶ 41:14
Insight
Unified Transformer architectures simplified the training of natively multimodal AI models
“Even if this is not, like, the only architecture that would be, like, in a multi-model, but it made it really simple to train these models, like, natively, because you have, like, a single architecture and you can have all the modalities in, during training.”
Mostafa Dehghani Apr 2, 2026 ▶ 43:32
Insight
Video data conveys physical world knowledge to AI more efficiently than text
“So because of that, like picking up a lot of knowledge about the word through language is just not really efficient. I don't want to say that it's impossible, but it's not efficient, you know, like to learn about gravity. If you kind of like, you know, have yo…”
Mostafa Dehghani Apr 2, 2026 ▶ 47:19
Insight
Demonstrating that image training lowers text perplexity remains extremely difficult
“So it turned out to be a really, really good model, but it was like really hard to see that. Wow. You know, I train on images and then like Text perplexity goes down. That was hard to see. You know, like the fact that, you know, you train in native model and i…”
Mostafa Dehghani Apr 2, 2026 ▶ 48:54
Insight
Interleaved multimodal generation overcomes single-shot limits via step-by-step planning
“But if you have incremental generation, so if you have text and then an image and text and image, you can get your model to generate these details one by one. So you never expect your model to generate an image, a perfect image in the first shot, right? So, so…”
Mostafa Dehghani Apr 2, 2026 ▶ 51:37
Insight
Jagged intelligence in AI is a structural flaw, not a patchable bug
“Not easy to pinpoint like specific things, but again, like, you know, this is just like my personal opinion and maybe I have colleagues and like the other people like sharing this with me, but I think we're underestimating how hard like jagged intelligence is …”
Mostafa Dehghani Apr 2, 2026 ▶ 54:59
Prediction Not checkable as stated
RAG will shift from universal use to handling long-tail distribution cases
“Maybe it changes in a way that, you know, like it doesn't need to trigger RAG for like everything, but I'm pretty sure that they're going to be some tail of the distribution that we're going to do RAG still for it.”
Mostafa Dehghani Apr 2, 2026 ▶ 58:09
Insight
Single AI failures damage user trust more than average performance builds it
“People don't experience average performance of these models. They experience, like, the failures. If you have your model doing a dumb mistake, the damage in the trust that it makes is, like, bigger than, like, you know, the benefit of getting hundred things ri…”
Mostafa Dehghani Apr 2, 2026 ▶ 1:01:46
Prediction Not checkable as stated
Real-world physical grounding will become the bottleneck for AI self-improvement
“As I said, you know, like soon, like the concept of data, like, you know, how to kind of like, you know, enable these models to kind of like, you know, be very good at like self-improvement becomes, how can I ground these models in, in real world? So this is d…”
Mostafa Dehghani Apr 2, 2026 ▶ 1:02:28
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.