Nov 14, 2024 · 39m · no-priors
No Priors Ep. 90 | With Google's DeepMind's AlphaProof Team
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
Google DeepMind's AlphaProof researchers discuss adapting AlphaZero's reinforcement learning framework to formal mathematical reasoning in Lean, detailing its breakthrough performance at the International Mathematical Olympiad, current architectural limitations, and transformative applications for research and software engineering.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 22.6% of the talking time here. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
After Sarah terms his point about human data 'pretty damning' and questions the relevance of human preference, Laurent immediately pushes back, clarifying that expert data fundamentally resolves exploration bottlenecks and saves years of compute.
Hardest push from the hosts ▶ 25:44 Sarah challenges the necessity of human interpretability over pure capabilitySarah directly challenges Laurent's premise that human mathematician data is primarily useful for proof aesthetics, arguing that capability advancements and alien proofs are what truly matter.
Biggest teaching moment ▶ 33:00 Rishi breaks down the alien solution to IMO Problem 6Rishi provides an intricate technical breakdown of the Aquasulian problem, demonstrating how AlphaProof invented an unconventional ceiling-function construction that human competitors and Fields Medalist Tim Gowers struggled to find.
The host holds their own ▶ 18:20 Elad demonstrates broad historical knowledge of applied pure mathematicsElad demonstrates domain expertise by synthesizing the historical pipeline of pure mathematics, citing group theory's role in quantum mechanics and number theory's application to zero-knowledge proofs in cryptography.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Architectural Foundation: Adapting AlphaZero to Formal Mathematical Proofs | 5 | 4 | 1 | 1 | Sarah sets the context with the International Mathematical Olympiad's prestige and asks how AlphaProof adapts game-playing search to mathematics. Thomas explains how AlphaZero's reinforcement learning and tree search were ported to infinite action spaces using the Lean formal language. | |
| Mathematical Search Spaces and Test-Time Reinforcement Learning | 5 | 6 | 1 | 1 | Sarah asks why some problems solved in minutes while others took three days. Rishi educates on how AlphaProof handles massive mathematical search spaces using 'test-time RL', generating problem variants to climb towards difficult proofs. | |
| Core Limitations: Theory Building and Formalizing Combinatorics | 5 | 6 | 1 | 1 | Elad inquires about scaling bottlenecks and problem domains AlphaProof cannot yet handle. Laurent and Thomas explain that AlphaProof cannot perform theory building or invent new mathematical objects, which currently hampers combinatorics formalization. | |
| Tackling Grand Mathematical Challenges and Millennium Prize Problems | 6 | 4 | 1 | 1 | Elad references Hilbert's 23 problems from 1900 to ask about grand AI milestones in math. Thomas discusses the Millennium Prize Problems, pointing out the unknown orders of magnitude separating current AI from conjectures like the Riemann hypothesis. | |
| Foundational Motivations: Math as a Testbed for Machine Reasoning and AGI | 6 | 3 | 1 | 1 | Sarah explains the Riemann hypothesis to listeners and asks whether the team views math as an end in itself or as a stepping stone to general reasoning. Each guest outlines their motivation, emphasizing math as a clean sandbox for scaling compute and pure truth-seeking. | |
| Real-World Applications: Formal Software Verification and Reasoning Transfer | 7 | 4 | 1 | 1 | Elad cites historical precedents like group theory in quantum physics and cryptography to ask about downstream practical applications. Laurent details formal software verification in mission-critical software, while Rishi discusses general reasoning transfer. | |
| Applying Reinforcement Learning Beyond Ground Truth and Expert Data Dynamics | 6 | 5 | 3 | 5 | Sarah asks how RL applies to subjective domains like humor and questions whether expert human labeling from figures like Terence Tao is even useful. When Laurent suggests human data merely guides 'niceness', Sarah pushes back that aesthetics matter less than capability, prompting Laurent and Rishi to clarify human data's role in bypassing exploration bottlenecks. | |
| Reimagining Mathematical Collaboration Through Automated Formal Verification | 6 | 5 | 1 | 1 | Elad brings up Poincaré as the last universal mathematician to ask how DeepMind engages the research community. Thomas clarifies that AlphaProof operates at a high-school competition level, but cites Terence Tao on how automated verification will allow trustless massive-scale mathematical collaboration. | |
| Case Study: AlphaProof's Alien Solution to IMO 2024 Problem 6 | 5 | 7 | 0 | 0 | Sarah asks about alien discoveries in AlphaProof's proofs similar to Move 37 in AlphaGo. Rishi provides a detailed visual walkthrough of IMO 2024 Problem 6, revealing how AlphaProof discovered a bizarre ceiling-function construction that stumped Fields Medalists. |