Nov 14, 2024 · 39m · no-priors

No Priors Ep. 90 | With Google's DeepMind's AlphaProof Team

Thomas Hubert · 13m spoken Rishi Mehta · 9m spoken Sarah Guo · 6m spoken Laurent Sartran · 5m spoken Elad Gil · 2m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Google DeepMind's AlphaProof researchers discuss adapting AlphaZero's reinforcement learning framework to formal mathematical reasoning in Lean, detailing its breakthrough performance at the International Mathematical Olympiad, current architectural limitations, and transformative applications for research and software engineering.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 22.6% of the talking time here. How this is scored →

The hosts as informed peer 5.7 Guest teaching 4.9 Guest disagreement 1.1 The hosts pushing back 1.3
05100:0010:0020:0030:002:49–6:27 · The hosts as informed peer 5/10 Architectural Foundation: Adapting AlphaZero to Formal Mathematical Proofs Sarah sets the context with the International Mathematical Olympiad's prestige and asks how AlphaProof adapts game-playing search to mathematics. Thomas explains how AlphaZero's reinforcement learning and tree search were ported to infinite action spaces using the Lean formal language.6:27–9:28 · The hosts as informed peer 5/10 Mathematical Search Spaces and Test-Time Reinforcement Learning Sarah asks why some problems solved in minutes while others took three days. Rishi educates on how AlphaProof handles massive mathematical search spaces using 'test-time RL', generating problem variants to climb towards difficult proofs.9:28–12:08 · The hosts as informed peer 5/10 Core Limitations: Theory Building and Formalizing Combinatorics Elad inquires about scaling bottlenecks and problem domains AlphaProof cannot yet handle. Laurent and Thomas explain that AlphaProof cannot perform theory building or invent new mathematical objects, which currently hampers combinatorics formalization.12:09–14:11 · The hosts as informed peer 6/10 Tackling Grand Mathematical Challenges and Millennium Prize Problems Elad references Hilbert's 23 problems from 1900 to ask about grand AI milestones in math. Thomas discusses the Millennium Prize Problems, pointing out the unknown orders of magnitude separating current AI from conjectures like the Riemann hypothesis.14:11–18:20 · The hosts as informed peer 6/10 Foundational Motivations: Math as a Testbed for Machine Reasoning and AGI Sarah explains the Riemann hypothesis to listeners and asks whether the team views math as an end in itself or as a stepping stone to general reasoning. Each guest outlines their motivation, emphasizing math as a clean sandbox for scaling compute and pure truth-seeking.18:20–22:00 · The hosts as informed peer 7/10 Real-World Applications: Formal Software Verification and Reasoning Transfer Elad cites historical precedents like group theory in quantum physics and cryptography to ask about downstream practical applications. Laurent details formal software verification in mission-critical software, while Rishi discusses general reasoning transfer.22:01–27:52 · The hosts as informed peer 6/10 Applying Reinforcement Learning Beyond Ground Truth and Expert Data Dynamics Sarah asks how RL applies to subjective domains like humor and questions whether expert human labeling from figures like Terence Tao is even useful. When Laurent suggests human data merely guides 'niceness', Sarah pushes back that aesthetics matter less than capability, prompting Laurent and Rishi to clarify human data's role in bypassing exploration bottlenecks.27:52–30:45 · The hosts as informed peer 6/10 Reimagining Mathematical Collaboration Through Automated Formal Verification Elad brings up Poincaré as the last universal mathematician to ask how DeepMind engages the research community. Thomas clarifies that AlphaProof operates at a high-school competition level, but cites Terence Tao on how automated verification will allow trustless massive-scale mathematical collaboration.30:45–34:45 · The hosts as informed peer 5/10 Case Study: AlphaProof's Alien Solution to IMO 2024 Problem 6 Sarah asks about alien discoveries in AlphaProof's proofs similar to Move 37 in AlphaGo. Rishi provides a detailed visual walkthrough of IMO 2024 Problem 6, revealing how AlphaProof discovered a bizarre ceiling-function construction that stumped Fields Medalists.2:49–6:27 · Guest teaching 4/10 Architectural Foundation: Adapting AlphaZero to Formal Mathematical Proofs Sarah sets the context with the International Mathematical Olympiad's prestige and asks how AlphaProof adapts game-playing search to mathematics. Thomas explains how AlphaZero's reinforcement learning and tree search were ported to infinite action spaces using the Lean formal language.6:27–9:28 · Guest teaching 6/10 Mathematical Search Spaces and Test-Time Reinforcement Learning Sarah asks why some problems solved in minutes while others took three days. Rishi educates on how AlphaProof handles massive mathematical search spaces using 'test-time RL', generating problem variants to climb towards difficult proofs.9:28–12:08 · Guest teaching 6/10 Core Limitations: Theory Building and Formalizing Combinatorics Elad inquires about scaling bottlenecks and problem domains AlphaProof cannot yet handle. Laurent and Thomas explain that AlphaProof cannot perform theory building or invent new mathematical objects, which currently hampers combinatorics formalization.12:09–14:11 · Guest teaching 4/10 Tackling Grand Mathematical Challenges and Millennium Prize Problems Elad references Hilbert's 23 problems from 1900 to ask about grand AI milestones in math. Thomas discusses the Millennium Prize Problems, pointing out the unknown orders of magnitude separating current AI from conjectures like the Riemann hypothesis.14:11–18:20 · Guest teaching 3/10 Foundational Motivations: Math as a Testbed for Machine Reasoning and AGI Sarah explains the Riemann hypothesis to listeners and asks whether the team views math as an end in itself or as a stepping stone to general reasoning. Each guest outlines their motivation, emphasizing math as a clean sandbox for scaling compute and pure truth-seeking.18:20–22:00 · Guest teaching 4/10 Real-World Applications: Formal Software Verification and Reasoning Transfer Elad cites historical precedents like group theory in quantum physics and cryptography to ask about downstream practical applications. Laurent details formal software verification in mission-critical software, while Rishi discusses general reasoning transfer.22:01–27:52 · Guest teaching 5/10 Applying Reinforcement Learning Beyond Ground Truth and Expert Data Dynamics Sarah asks how RL applies to subjective domains like humor and questions whether expert human labeling from figures like Terence Tao is even useful. When Laurent suggests human data merely guides 'niceness', Sarah pushes back that aesthetics matter less than capability, prompting Laurent and Rishi to clarify human data's role in bypassing exploration bottlenecks.27:52–30:45 · Guest teaching 5/10 Reimagining Mathematical Collaboration Through Automated Formal Verification Elad brings up Poincaré as the last universal mathematician to ask how DeepMind engages the research community. Thomas clarifies that AlphaProof operates at a high-school competition level, but cites Terence Tao on how automated verification will allow trustless massive-scale mathematical collaboration.30:45–34:45 · Guest teaching 7/10 Case Study: AlphaProof's Alien Solution to IMO 2024 Problem 6 Sarah asks about alien discoveries in AlphaProof's proofs similar to Move 37 in AlphaGo. Rishi provides a detailed visual walkthrough of IMO 2024 Problem 6, revealing how AlphaProof discovered a bizarre ceiling-function construction that stumped Fields Medalists.2:49–6:27 · Guest disagreement 1/10 Architectural Foundation: Adapting AlphaZero to Formal Mathematical Proofs Sarah sets the context with the International Mathematical Olympiad's prestige and asks how AlphaProof adapts game-playing search to mathematics. Thomas explains how AlphaZero's reinforcement learning and tree search were ported to infinite action spaces using the Lean formal language.6:27–9:28 · Guest disagreement 1/10 Mathematical Search Spaces and Test-Time Reinforcement Learning Sarah asks why some problems solved in minutes while others took three days. Rishi educates on how AlphaProof handles massive mathematical search spaces using 'test-time RL', generating problem variants to climb towards difficult proofs.9:28–12:08 · Guest disagreement 1/10 Core Limitations: Theory Building and Formalizing Combinatorics Elad inquires about scaling bottlenecks and problem domains AlphaProof cannot yet handle. Laurent and Thomas explain that AlphaProof cannot perform theory building or invent new mathematical objects, which currently hampers combinatorics formalization.12:09–14:11 · Guest disagreement 1/10 Tackling Grand Mathematical Challenges and Millennium Prize Problems Elad references Hilbert's 23 problems from 1900 to ask about grand AI milestones in math. Thomas discusses the Millennium Prize Problems, pointing out the unknown orders of magnitude separating current AI from conjectures like the Riemann hypothesis.14:11–18:20 · Guest disagreement 1/10 Foundational Motivations: Math as a Testbed for Machine Reasoning and AGI Sarah explains the Riemann hypothesis to listeners and asks whether the team views math as an end in itself or as a stepping stone to general reasoning. Each guest outlines their motivation, emphasizing math as a clean sandbox for scaling compute and pure truth-seeking.18:20–22:00 · Guest disagreement 1/10 Real-World Applications: Formal Software Verification and Reasoning Transfer Elad cites historical precedents like group theory in quantum physics and cryptography to ask about downstream practical applications. Laurent details formal software verification in mission-critical software, while Rishi discusses general reasoning transfer.22:01–27:52 · Guest disagreement 3/10 Applying Reinforcement Learning Beyond Ground Truth and Expert Data Dynamics Sarah asks how RL applies to subjective domains like humor and questions whether expert human labeling from figures like Terence Tao is even useful. When Laurent suggests human data merely guides 'niceness', Sarah pushes back that aesthetics matter less than capability, prompting Laurent and Rishi to clarify human data's role in bypassing exploration bottlenecks.27:52–30:45 · Guest disagreement 1/10 Reimagining Mathematical Collaboration Through Automated Formal Verification Elad brings up Poincaré as the last universal mathematician to ask how DeepMind engages the research community. Thomas clarifies that AlphaProof operates at a high-school competition level, but cites Terence Tao on how automated verification will allow trustless massive-scale mathematical collaboration.30:45–34:45 · Guest disagreement 0/10 Case Study: AlphaProof's Alien Solution to IMO 2024 Problem 6 Sarah asks about alien discoveries in AlphaProof's proofs similar to Move 37 in AlphaGo. Rishi provides a detailed visual walkthrough of IMO 2024 Problem 6, revealing how AlphaProof discovered a bizarre ceiling-function construction that stumped Fields Medalists.2:49–6:27 · The hosts pushing back 1/10 Architectural Foundation: Adapting AlphaZero to Formal Mathematical Proofs Sarah sets the context with the International Mathematical Olympiad's prestige and asks how AlphaProof adapts game-playing search to mathematics. Thomas explains how AlphaZero's reinforcement learning and tree search were ported to infinite action spaces using the Lean formal language.6:27–9:28 · The hosts pushing back 1/10 Mathematical Search Spaces and Test-Time Reinforcement Learning Sarah asks why some problems solved in minutes while others took three days. Rishi educates on how AlphaProof handles massive mathematical search spaces using 'test-time RL', generating problem variants to climb towards difficult proofs.9:28–12:08 · The hosts pushing back 1/10 Core Limitations: Theory Building and Formalizing Combinatorics Elad inquires about scaling bottlenecks and problem domains AlphaProof cannot yet handle. Laurent and Thomas explain that AlphaProof cannot perform theory building or invent new mathematical objects, which currently hampers combinatorics formalization.12:09–14:11 · The hosts pushing back 1/10 Tackling Grand Mathematical Challenges and Millennium Prize Problems Elad references Hilbert's 23 problems from 1900 to ask about grand AI milestones in math. Thomas discusses the Millennium Prize Problems, pointing out the unknown orders of magnitude separating current AI from conjectures like the Riemann hypothesis.14:11–18:20 · The hosts pushing back 1/10 Foundational Motivations: Math as a Testbed for Machine Reasoning and AGI Sarah explains the Riemann hypothesis to listeners and asks whether the team views math as an end in itself or as a stepping stone to general reasoning. Each guest outlines their motivation, emphasizing math as a clean sandbox for scaling compute and pure truth-seeking.18:20–22:00 · The hosts pushing back 1/10 Real-World Applications: Formal Software Verification and Reasoning Transfer Elad cites historical precedents like group theory in quantum physics and cryptography to ask about downstream practical applications. Laurent details formal software verification in mission-critical software, while Rishi discusses general reasoning transfer.22:01–27:52 · The hosts pushing back 5/10 Applying Reinforcement Learning Beyond Ground Truth and Expert Data Dynamics Sarah asks how RL applies to subjective domains like humor and questions whether expert human labeling from figures like Terence Tao is even useful. When Laurent suggests human data merely guides 'niceness', Sarah pushes back that aesthetics matter less than capability, prompting Laurent and Rishi to clarify human data's role in bypassing exploration bottlenecks.27:52–30:45 · The hosts pushing back 1/10 Reimagining Mathematical Collaboration Through Automated Formal Verification Elad brings up Poincaré as the last universal mathematician to ask how DeepMind engages the research community. Thomas clarifies that AlphaProof operates at a high-school competition level, but cites Terence Tao on how automated verification will allow trustless massive-scale mathematical collaboration.30:45–34:45 · The hosts pushing back 0/10 Case Study: AlphaProof's Alien Solution to IMO 2024 Problem 6 Sarah asks about alien discoveries in AlphaProof's proofs similar to Move 37 in AlphaGo. Rishi provides a detailed visual walkthrough of IMO 2024 Problem 6, revealing how AlphaProof discovered a bizarre ceiling-function construction that stumped Fields Medalists.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 28.3% · guest 71.7%0:00 · the hosts 28.3% · guest 71.7%3:00 · the hosts 20.5% · guest 79.5%3:00 · the hosts 20.5% · guest 79.5%6:00 · the hosts 11.2% · guest 88.8%6:00 · the hosts 11.2% · guest 88.8%9:00 · the hosts 19.8% · guest 80.2%9:00 · the hosts 19.8% · guest 80.2%12:00 · the hosts 36.3% · guest 63.7%12:00 · the hosts 36.3% · guest 63.7%15:00 · the hosts 6.4% · guest 93.6%15:00 · the hosts 6.4% · guest 93.6%18:00 · the hosts 38.6% · guest 61.4%18:00 · the hosts 38.6% · guest 61.4%21:00 · the hosts 24.9% · guest 75.1%21:00 · the hosts 24.9% · guest 75.1%24:00 · the hosts 25.8% · guest 74.2%24:00 · the hosts 25.8% · guest 74.2%27:00 · the hosts 21.6% · guest 78.4%27:00 · the hosts 21.6% · guest 78.4%30:00 · the hosts 11.1% · guest 88.9%30:00 · the hosts 11.1% · guest 88.9%33:00 · the hosts 5.4% · guest 94.6%33:00 · the hosts 5.4% · guest 94.6%36:00 · the hosts 36.1% · guest 63.9%36:00 · the hosts 36.1% · guest 63.9%39:00 · the hosts 100% · guest 0%39:00 · the hosts 100% · guest 0%
Sharpest disagreement ▶ 26:02 Laurent counters Sarah's dismissive take on expert human data

After Sarah terms his point about human data 'pretty damning' and questions the relevance of human preference, Laurent immediately pushes back, clarifying that expert data fundamentally resolves exploration bottlenecks and saves years of compute.

Hardest push from the hosts ▶ 25:44 Sarah challenges the necessity of human interpretability over pure capability

Sarah directly challenges Laurent's premise that human mathematician data is primarily useful for proof aesthetics, arguing that capability advancements and alien proofs are what truly matter.

Biggest teaching moment ▶ 33:00 Rishi breaks down the alien solution to IMO Problem 6

Rishi provides an intricate technical breakdown of the Aquasulian problem, demonstrating how AlphaProof invented an unconventional ceiling-function construction that human competitors and Fields Medalist Tim Gowers struggled to find.

The host holds their own ▶ 18:20 Elad demonstrates broad historical knowledge of applied pure mathematics

Elad demonstrates domain expertise by synthesizing the historical pipeline of pure mathematics, citing group theory's role in quantum mechanics and number theory's application to zero-knowledge proofs in cryptography.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Architectural Foundation: Adapting AlphaZero to Formal Mathematical Proofs 5411 Sarah sets the context with the International Mathematical Olympiad's prestige and asks how AlphaProof adapts game-playing search to mathematics. Thomas explains how AlphaZero's reinforcement learning and tree search were ported to infinite action spaces using the Lean formal language.
Mathematical Search Spaces and Test-Time Reinforcement Learning 5611 Sarah asks why some problems solved in minutes while others took three days. Rishi educates on how AlphaProof handles massive mathematical search spaces using 'test-time RL', generating problem variants to climb towards difficult proofs.
Core Limitations: Theory Building and Formalizing Combinatorics 5611 Elad inquires about scaling bottlenecks and problem domains AlphaProof cannot yet handle. Laurent and Thomas explain that AlphaProof cannot perform theory building or invent new mathematical objects, which currently hampers combinatorics formalization.
Tackling Grand Mathematical Challenges and Millennium Prize Problems 6411 Elad references Hilbert's 23 problems from 1900 to ask about grand AI milestones in math. Thomas discusses the Millennium Prize Problems, pointing out the unknown orders of magnitude separating current AI from conjectures like the Riemann hypothesis.
Foundational Motivations: Math as a Testbed for Machine Reasoning and AGI 6311 Sarah explains the Riemann hypothesis to listeners and asks whether the team views math as an end in itself or as a stepping stone to general reasoning. Each guest outlines their motivation, emphasizing math as a clean sandbox for scaling compute and pure truth-seeking.
Real-World Applications: Formal Software Verification and Reasoning Transfer 7411 Elad cites historical precedents like group theory in quantum physics and cryptography to ask about downstream practical applications. Laurent details formal software verification in mission-critical software, while Rishi discusses general reasoning transfer.
Applying Reinforcement Learning Beyond Ground Truth and Expert Data Dynamics 6535 Sarah asks how RL applies to subjective domains like humor and questions whether expert human labeling from figures like Terence Tao is even useful. When Laurent suggests human data merely guides 'niceness', Sarah pushes back that aesthetics matter less than capability, prompting Laurent and Rishi to clarify human data's role in bypassing exploration bottlenecks.
Reimagining Mathematical Collaboration Through Automated Formal Verification 6511 Elad brings up Poincaré as the last universal mathematician to ask how DeepMind engages the research community. Thomas clarifies that AlphaProof operates at a high-school competition level, but cites Terence Tao on how automated verification will allow trustless massive-scale mathematical collaboration.
Case Study: AlphaProof's Alien Solution to IMO 2024 Problem 6 5700 Sarah asks about alien discoveries in AlphaProof's proofs similar to Move 37 in AlphaGo. Rishi provides a detailed visual walkthrough of IMO 2024 Problem 6, revealing how AlphaProof discovered a bizarre ceiling-function construction that stumped Fields Medalists.

Statements from this episode (16)

Assertion Partly supported
Guo: DeepMind's AlphaProof solved four of six 2024 IMO problems
“Alphaproof had this really amazing results of solving four of the six problems this year.”
Sarah Guo Nov 14, 2024 ▶ 3:05
Insight
Hubert: Formal math proofs enable self-improving reinforcement learning loops
“The advantage of that is that once kind of the proof is complete then you know, the machine would give you a signal back to say, yes, your proof is correct or not. And so we could search for kind of correct proofs. Once we find a correct proof, we can learn fr…”
Thomas Hubert Nov 14, 2024 ▶ 5:55
Assertion Supported
Mehta: AlphaProof Is Strongest in IMO Algebra and Number Theory
“So the IMO problems have come in four categories. So there's algebra number theory, combinatrix, and geometry. The two that it's strongest at are algebra and number theory. It's relatively weaker at combinatrix, although it can do quite good at some IMO combin…”
Rishi Mehta Nov 14, 2024 ▶ 7:51
Insight
Mehta: AlphaProof Uses Test-Time RL on Problem Variations to Find Proofs
“One of the ways in which it navigates this massive search space is via an idea that we came up with, which we call test time RL. So this is an idea where, like, let's say you're confronted with a problem that you don't know how to solve. And you know, you can …”
Rishi Mehta Nov 14, 2024 ▶ 8:13
Assertion Supported
Sartran: AlphaProof cannot perform theory building or invent new theories
“Maybe the main thing that Alphaproof doesn't do is theory building. It doesn't invent its own theories.”
Laurent Sartran Nov 14, 2024 ▶ 10:04
Insight
DeepMind's Hubert: Solving Harder Math Requires Introducing New Mathematical Objects
“To prove harder problems, you start to need to be able to introduce these new kind of mathematical objects to decompose the problems or problems. And we see that already happening for the IMO at a, you know, small scale.”
Thomas Hubert Nov 14, 2024 ▶ 13:46
Assertion Supported
Sartran: Formal code verification remains bottlenecked by human proof-writing
“So it's already done for very critical domain, like avionics cryptography where it's very important, but it has to be done by humans.”
Laurent Sartran Nov 14, 2024 ▶ 21:00
Opinion
Mehta: AlphaProof's RL scaling and test-time compute generalize across domains
“Some of the sort of tech we developed here of like, you know, like scaling RL and like figuring out how to spend a lot of inference time compute stuff like this feels like it's Quite generally applicable to many other problems.”
Rishi Mehta Nov 14, 2024 ▶ 21:49
Disclosure
Sartran: AlphaProof develops an alien style by learning from self-discovered proofs
“The way alpha proof currency operates there is that it discovers its own proofs, and when they are valid, it learns from them and develop its own style which has been commented upon as looking, yeah, quite, quite alien.”
Laurent Sartran Nov 14, 2024 ▶ 24:57
Assertion Not checkable as stated
Sartran: Translating all known human proofs would save years of compute
“With more supervised data, we can avoid the exploration problem, and we could translate all the proofs that are known to man, and that certainly would make the agent much better, and that would save us years of compute for sure.”
Laurent Sartran Nov 14, 2024 ▶ 26:07
Prediction Not checkable as stated
Mehta: Specialist data seeds LLMs, but RL drives superhuman capability
“The specialist humans are gonna serve to, like, get the LLM from, like, just a bunch of weights that knows how to do nothing to, like, something that, like is, like, surprisingly strong. And then the RL is gonna take you from there to, like, something that's, …”
Rishi Mehta Nov 14, 2024 ▶ 27:32
Assertion Supported
Hubert: AlphaProof reached high school level but cannot rival Terence Tao
“And to be honest with you know, like kind of we, at the moment we can't rival at all with someone like Terry Tao. We, I think we demonstrated that what we've demonstrated is that we can learn general mathematics almost from scratch and arrive at kind of an imp…”
Thomas Hubert Nov 14, 2024 ▶ 28:36
Insight
Hubert: Automated Formal Proof Checkers Enable Mass Crowdsourced Mathematics
“But if you instead relied on a formal system to check everyone else's work, then You could do a little bit like in astronomy where you could have an amateur kind of living in the middle of maybe nowhere and you, you've never met. And then you wouldn't have to …”
Thomas Hubert Nov 14, 2024 ▶ 29:24
Assertion Not checkable as stated
Mehta: Fields Medalist Tim Gowers failed to find AlphaProof's IMO construction
“Tim Gowers, who was one of our judges and is also a fields medalist. Tried this question for a couple hours and he couldn't find the construction for function that had this property”
Rishi Mehta Nov 14, 2024 ▶ 33:55
Insight
Sartran: Human creativity matters more in theory building than theorem proving
“It seems that there is more room for human creativity, human taste human skill in building theories than in the proof part.”
Laurent Sartran Nov 14, 2024 ▶ 37:04
Insight
Mehta: As AI solves problems, human roles shift toward question framing
“As machines get better at finding the answers, like, we're going to have to get better at finding the questions. And, you know, like these systems don't have a, you know, their own sort of notion of what questions are interesting. And given a large question, h…”
Rishi Mehta Nov 14, 2024 ▶ 38:19
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.