Jun 26, 2025 · 42m · no-priors

No Priors Ep. 120 | With Google DeepMind’s Pushmeet Kohli and Matej Balog

Pushmeet Kohli · 17m spoken Matej Balog · 15m spoken Sarah Guo · 6m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of No Priors, Google DeepMind researchers Pushmeet Kohli and Matej Balog explore Alpha Evolve, an autonomous AI coding agent combining large language models and evolutionary search to discover novel, superhuman algorithms. They detail the system's architecture, successful infrastructure deployments across Google, and the transformative potential of human-AI collaboration in scientific research.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 15.4% of the talking time here. How this is scored →

The hosts as informed peer 6.0 Guest teaching 4.7 Guest disagreement 1.4 The hosts pushing back 2.3
05100:0015:0030:001:16–8:02 · The hosts as informed peer 5/10 Lineage of Algorithmic Discovery from AlphaGo to FunSearch Sarah sets the context and asks how Alpha Evolve differentiates itself from predecessors like AlphaTensor and FunSearch. Pushmeet provides an extensive technical masterclass tracing the lineage back to AlphaGo's Move 37, Strassen's matrix multiplication, and program-space search.8:02–11:18 · The hosts as informed peer 6/10 Unpacking Technical Creativity and Massive Algorithmic Search Spaces Sarah asks whether past algorithms went undiscovered due to vast search spaces or human complacency. Matej gently but firmly pushes back on the complacency premise, explaining that these benchmark problems have been aggressively pursued by world-class researchers for decades.11:18–16:52 · The hosts as informed peer 6/10 Mechanics of Alpha Evolve and Test-Time Scaling Sarah requests a concrete breakdown of data center scheduling optimization and queries test-time scaling dynamics. Matej walks through evaluation functions and evolutionary populations, explaining how difficulty dictates inference runtime.16:52–25:21 · The hosts as informed peer 7/10 Harnessing LLM Hallucinations Through Rigorous Automated Evaluators Sarah probes how developer coding agents fail and how automated evaluators can overcome specification bottlenecks. Pushmeet explains harnessing LLM hallucinations via rigorous evaluators, and Matej outlines the continuum between hard simulators and LLM-as-a-judge critiques.25:22–31:50 · The hosts as informed peer 7/10 Evaluating Recursive Self-Improvement and Broader Scientific Applications Sarah presses on whether Alpha Evolve's 23% speedup of Gemini's training stack constitutes genuine recursive self-improvement. Pushmeet and Matej qualify the milestone, noting it improves compute efficiency rather than core cognitive capability, and map out mathematical extensions.31:50–38:29 · The hosts as informed peer 6/10 Human-AI Collaboration and Code Interpretability in Science Sarah asks about human scientist roles when physical lab automation converges with AI search. Pushmeet and Matej describe human-AI collaboration, highlighting that discovering interpretable algorithmic code is vastly superior to black-box neural policies.38:29–42:05 · The hosts as informed peer 5/10 Broader Access, Infrastructure Deployment, and Conclusion Sarah closes by asking about broad accessibility and unannounced internal Google deployments. Pushmeet outlines the trusted tester initiative and compute requirements, while Matej summarizes full-stack infrastructure applications.1:16–8:02 · Guest teaching 6/10 Lineage of Algorithmic Discovery from AlphaGo to FunSearch Sarah sets the context and asks how Alpha Evolve differentiates itself from predecessors like AlphaTensor and FunSearch. Pushmeet provides an extensive technical masterclass tracing the lineage back to AlphaGo's Move 37, Strassen's matrix multiplication, and program-space search.8:02–11:18 · Guest teaching 5/10 Unpacking Technical Creativity and Massive Algorithmic Search Spaces Sarah asks whether past algorithms went undiscovered due to vast search spaces or human complacency. Matej gently but firmly pushes back on the complacency premise, explaining that these benchmark problems have been aggressively pursued by world-class researchers for decades.11:18–16:52 · Guest teaching 5/10 Mechanics of Alpha Evolve and Test-Time Scaling Sarah requests a concrete breakdown of data center scheduling optimization and queries test-time scaling dynamics. Matej walks through evaluation functions and evolutionary populations, explaining how difficulty dictates inference runtime.16:52–25:21 · Guest teaching 5/10 Harnessing LLM Hallucinations Through Rigorous Automated Evaluators Sarah probes how developer coding agents fail and how automated evaluators can overcome specification bottlenecks. Pushmeet explains harnessing LLM hallucinations via rigorous evaluators, and Matej outlines the continuum between hard simulators and LLM-as-a-judge critiques.25:22–31:50 · Guest teaching 4/10 Evaluating Recursive Self-Improvement and Broader Scientific Applications Sarah presses on whether Alpha Evolve's 23% speedup of Gemini's training stack constitutes genuine recursive self-improvement. Pushmeet and Matej qualify the milestone, noting it improves compute efficiency rather than core cognitive capability, and map out mathematical extensions.31:50–38:29 · Guest teaching 5/10 Human-AI Collaboration and Code Interpretability in Science Sarah asks about human scientist roles when physical lab automation converges with AI search. Pushmeet and Matej describe human-AI collaboration, highlighting that discovering interpretable algorithmic code is vastly superior to black-box neural policies.38:29–42:05 · Guest teaching 3/10 Broader Access, Infrastructure Deployment, and Conclusion Sarah closes by asking about broad accessibility and unannounced internal Google deployments. Pushmeet outlines the trusted tester initiative and compute requirements, while Matej summarizes full-stack infrastructure applications.1:16–8:02 · Guest disagreement 1/10 Lineage of Algorithmic Discovery from AlphaGo to FunSearch Sarah sets the context and asks how Alpha Evolve differentiates itself from predecessors like AlphaTensor and FunSearch. Pushmeet provides an extensive technical masterclass tracing the lineage back to AlphaGo's Move 37, Strassen's matrix multiplication, and program-space search.8:02–11:18 · Guest disagreement 3/10 Unpacking Technical Creativity and Massive Algorithmic Search Spaces Sarah asks whether past algorithms went undiscovered due to vast search spaces or human complacency. Matej gently but firmly pushes back on the complacency premise, explaining that these benchmark problems have been aggressively pursued by world-class researchers for decades.11:18–16:52 · Guest disagreement 1/10 Mechanics of Alpha Evolve and Test-Time Scaling Sarah requests a concrete breakdown of data center scheduling optimization and queries test-time scaling dynamics. Matej walks through evaluation functions and evolutionary populations, explaining how difficulty dictates inference runtime.16:52–25:21 · Guest disagreement 1/10 Harnessing LLM Hallucinations Through Rigorous Automated Evaluators Sarah probes how developer coding agents fail and how automated evaluators can overcome specification bottlenecks. Pushmeet explains harnessing LLM hallucinations via rigorous evaluators, and Matej outlines the continuum between hard simulators and LLM-as-a-judge critiques.25:22–31:50 · Guest disagreement 2/10 Evaluating Recursive Self-Improvement and Broader Scientific Applications Sarah presses on whether Alpha Evolve's 23% speedup of Gemini's training stack constitutes genuine recursive self-improvement. Pushmeet and Matej qualify the milestone, noting it improves compute efficiency rather than core cognitive capability, and map out mathematical extensions.31:50–38:29 · Guest disagreement 1/10 Human-AI Collaboration and Code Interpretability in Science Sarah asks about human scientist roles when physical lab automation converges with AI search. Pushmeet and Matej describe human-AI collaboration, highlighting that discovering interpretable algorithmic code is vastly superior to black-box neural policies.38:29–42:05 · Guest disagreement 1/10 Broader Access, Infrastructure Deployment, and Conclusion Sarah closes by asking about broad accessibility and unannounced internal Google deployments. Pushmeet outlines the trusted tester initiative and compute requirements, while Matej summarizes full-stack infrastructure applications.1:16–8:02 · The hosts pushing back 1/10 Lineage of Algorithmic Discovery from AlphaGo to FunSearch Sarah sets the context and asks how Alpha Evolve differentiates itself from predecessors like AlphaTensor and FunSearch. Pushmeet provides an extensive technical masterclass tracing the lineage back to AlphaGo's Move 37, Strassen's matrix multiplication, and program-space search.8:02–11:18 · The hosts pushing back 3/10 Unpacking Technical Creativity and Massive Algorithmic Search Spaces Sarah asks whether past algorithms went undiscovered due to vast search spaces or human complacency. Matej gently but firmly pushes back on the complacency premise, explaining that these benchmark problems have been aggressively pursued by world-class researchers for decades.11:18–16:52 · The hosts pushing back 2/10 Mechanics of Alpha Evolve and Test-Time Scaling Sarah requests a concrete breakdown of data center scheduling optimization and queries test-time scaling dynamics. Matej walks through evaluation functions and evolutionary populations, explaining how difficulty dictates inference runtime.16:52–25:21 · The hosts pushing back 3/10 Harnessing LLM Hallucinations Through Rigorous Automated Evaluators Sarah probes how developer coding agents fail and how automated evaluators can overcome specification bottlenecks. Pushmeet explains harnessing LLM hallucinations via rigorous evaluators, and Matej outlines the continuum between hard simulators and LLM-as-a-judge critiques.25:22–31:50 · The hosts pushing back 4/10 Evaluating Recursive Self-Improvement and Broader Scientific Applications Sarah presses on whether Alpha Evolve's 23% speedup of Gemini's training stack constitutes genuine recursive self-improvement. Pushmeet and Matej qualify the milestone, noting it improves compute efficiency rather than core cognitive capability, and map out mathematical extensions.31:50–38:29 · The hosts pushing back 2/10 Human-AI Collaboration and Code Interpretability in Science Sarah asks about human scientist roles when physical lab automation converges with AI search. Pushmeet and Matej describe human-AI collaboration, highlighting that discovering interpretable algorithmic code is vastly superior to black-box neural policies.38:29–42:05 · The hosts pushing back 1/10 Broader Access, Infrastructure Deployment, and Conclusion Sarah closes by asking about broad accessibility and unannounced internal Google deployments. Pushmeet outlines the trusted tester initiative and compute requirements, while Matej summarizes full-stack infrastructure applications.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 35.9% · guest 64.1%0:00 · the hosts 35.9% · guest 64.1%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 31.9% · guest 68.1%6:00 · the hosts 31.9% · guest 68.1%9:00 · the hosts 19% · guest 81%9:00 · the hosts 19% · guest 81%12:00 · the hosts 11% · guest 89%12:00 · the hosts 11% · guest 89%15:00 · the hosts 16.2% · guest 83.8%15:00 · the hosts 16.2% · guest 83.8%18:00 · the hosts 20.7% · guest 79.3%18:00 · the hosts 20.7% · guest 79.3%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 18% · guest 82%24:00 · the hosts 18% · guest 82%27:00 · the hosts 16.9% · guest 83.1%27:00 · the hosts 16.9% · guest 83.1%30:00 · the hosts 22.3% · guest 77.7%30:00 · the hosts 22.3% · guest 77.7%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 7.7% · guest 92.3%36:00 · the hosts 7.7% · guest 92.3%39:00 · the hosts 13.7% · guest 86.3%39:00 · the hosts 13.7% · guest 86.3%42:00 · the hosts 100% · guest 0%42:00 · the hosts 100% · guest 0%
Sharpest disagreement ▶ 9:50 Refuting the complacency hypothesis

Matej explicitly rejects Sarah's suggestion that prior algorithmic plateaus were caused by human complacency, emphasizing that these specific problems had top minds working intensely on them for decades.

Hardest push from the hosts ▶ 25:21 Pressing on recursive self-improvement

Sarah directly challenges the guests on whether demonstrated training infrastructure speedups mean we are witnessing the onset of recursive self-improvement, following up with whether there is reason to doubt it.

Biggest teaching moment ▶ 4:45 Strassen matrix multiplication lineage

Pushmeet delivers a thorough breakdown of Strassen's counterintuitive sub-cubic matrix multiplication complexity, explaining how AlphaTensor and FunSearch surpassed 50-year-old human-designed algorithmic baselines.

The host holds their own ▶ 19:42 Deconstructing evaluator bottlenecks and PM workflows

Sarah demonstrates sharp domain mastery by dissecting prompt specification limits, proposing natural language evaluators, domain simulators, and execution traces as solutions.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Lineage of Algorithmic Discovery from AlphaGo to FunSearch 5611 Sarah sets the context and asks how Alpha Evolve differentiates itself from predecessors like AlphaTensor and FunSearch. Pushmeet provides an extensive technical masterclass tracing the lineage back to AlphaGo's Move 37, Strassen's matrix multiplication, and program-space search.
Unpacking Technical Creativity and Massive Algorithmic Search Spaces 6533 Sarah asks whether past algorithms went undiscovered due to vast search spaces or human complacency. Matej gently but firmly pushes back on the complacency premise, explaining that these benchmark problems have been aggressively pursued by world-class researchers for decades.
Mechanics of Alpha Evolve and Test-Time Scaling 6512 Sarah requests a concrete breakdown of data center scheduling optimization and queries test-time scaling dynamics. Matej walks through evaluation functions and evolutionary populations, explaining how difficulty dictates inference runtime.
Harnessing LLM Hallucinations Through Rigorous Automated Evaluators 7513 Sarah probes how developer coding agents fail and how automated evaluators can overcome specification bottlenecks. Pushmeet explains harnessing LLM hallucinations via rigorous evaluators, and Matej outlines the continuum between hard simulators and LLM-as-a-judge critiques.
Evaluating Recursive Self-Improvement and Broader Scientific Applications 7424 Sarah presses on whether Alpha Evolve's 23% speedup of Gemini's training stack constitutes genuine recursive self-improvement. Pushmeet and Matej qualify the milestone, noting it improves compute efficiency rather than core cognitive capability, and map out mathematical extensions.
Human-AI Collaboration and Code Interpretability in Science 6512 Sarah asks about human scientist roles when physical lab automation converges with AI search. Pushmeet and Matej describe human-AI collaboration, highlighting that discovering interpretable algorithmic code is vastly superior to black-box neural policies.
Broader Access, Infrastructure Deployment, and Conclusion 5311 Sarah closes by asking about broad accessibility and unannounced internal Google deployments. Pushmeet outlines the trusted tester initiative and compute requirements, while Matej summarizes full-stack infrastructure applications.

Statements from this episode (14)

Assertion Supported
Balog: AlphaEvolve algorithms are already deployed in Google's infrastructure
“Alpha Evolve is an AI coding agent that is able to discover new algorithms that are able to make new discoveries on open scientific problems. And at the same time, those algorithms can be so practical that they are already deployed in keep Key parts of Google'…”
Matej Balog Jun 26, 2025 ▶ 0:57
Assertion Supported
Balog: AlphaTensor was the first AI to discover superhuman matrix multiplication algorithms
“And kind of the first breakthrough we had in this space was in twenty-twenty-two when we released a system called Alpha Tensor. And so that was a system that was an AI system using reinforcement learning. That for a very specific, but a fundamental computation…”
Matej Balog Jun 26, 2025 ▶ 1:52
Assertion Supported
Kohli: FunSearch made the first scientific discovery generated by an LLM
“FunSearch, which was an LLM-based agent, which for the first time, by searching in the space of programs, showed that you can come up with completely new solutions, like, and this, and made the first scientific discovery from an LLM.”
Pushmeet Kohli Jun 26, 2025 ▶ 7:33
Insight
Kohli: Large matrix multiplication algorithms are too intricate for human intuition alone
“As you sort of go to larger sizes, the space is so huge. The constructions are not sort of something which is very natural. These are very involved and intricate, intricate, intricate sort of constructions that would be very hard to discover by chance. So it's…”
Pushmeet Kohli Jun 26, 2025 ▶ 9:18
Disclosure
Balog: DeepMind benchmarked AlphaEvolve against heavily optimized Google infrastructure problems
“The problems we chose to apply Alpha Evolve to in the first instance both on the scientific side and the practical side, We deliberately chose problems which have been worked on for a very long time by the very best people. So on the, you know, on the scientif…”
Matej Balog Jun 26, 2025 ▶ 10:04
Disclosure
DeepMind: AlphaEvolve was seeded with human baselines for data center scheduling
“Another option you can take is actually we have already worked on this problem for a really long time. Here is a very strong initial solution that we can provide to the system and you can start from here. And that's what we did for the application to discoveri…”
Matej Balog Jun 26, 2025 ▶ 13:07
Assertion Supported
Kohli: Multi-agent LLMs outperform base models in hypothesis generation
“And we have shown that this AI co-scientist which we have used for hypothesis generation, it basically uses a multi-agent sort of setup and where LLMs themselves are able to sort of figure out that certain hypotheses are better in terms of novelty and signific…”
Pushmeet Kohli Jun 26, 2025 ▶ 24:37
Assertion Supported
Kohli: AlphaEvolve has successfully made AI model training more computationally efficient
“What Alpha Evolve has been able to do is basically make training more efficient.”
Pushmeet Kohli Jun 26, 2025 ▶ 26:01
Prediction Not checkable as stated
Kohli: AI self-improvement for cognitive tasks will work with proper evaluators
“It should work. But as we sort of mentioned that having good evaluators is an important element, right? And so having a sort of evaluator, which can say this proposal that you have just suggested for me to improve the training process will yield A good result.…”
Pushmeet Kohli Jun 26, 2025 ▶ 26:40
Disclosure
Balog: AlphaEvolve focuses on math and CS due to automated evaluators
“In Alpha Evolve, we've Primarily focused on mathematics and computer science, because these are the domains where it's the easiest to get these automated evaluation functions. Like you often get them basically for free.”
Matej Balog Jun 26, 2025 ▶ 29:09
Insight
Balog: Inspectable AI-generated code is safer than black-box neural networks
“It's hugely valuable that the artifact you get out of Alpha Evolve is a piece of code, and then you deploy that piece of code. And so before you do that experts, engineers who have worked on that system can visually inspect that piece of code, understand it, a…”
Matej Balog Jun 26, 2025 ▶ 36:51
Assertion Supported
Kohli: AlphaEvolve discovered cap set problem symmetries unknown to mathematicians
“Working with Jordan Ellenberg in the first, in the earlier version of Alpha Evolve, when you were working on the cap set problem, the programs that it discovered had very interesting symmetries that that mathematicians did not know about. And so, so not only t…”
Pushmeet Kohli Jun 26, 2025 ▶ 37:59
Disclosure
Kohli: DeepMind launches trusted tester program for AlphaEvolve
“We have started a trusted tester program where we have asked people to submit proposals, and what we intend to do with that program is to figure out what are the right ways in which people can really leverage alpha evolved.”
Pushmeet Kohli Jun 26, 2025 ▶ 38:53
Assertion Supported
Balog: AlphaEvolve improves Google data center efficiency, hardware design, and software
“So, so we show that AlphaEvolve can improve the efficiency of the data center. It can contribute to hardware design, and it can contribute to improving the efficiency of most important pieces of software that that are being run inside Google.”
Matej Balog Jun 26, 2025 ▶ 40:49
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.