Feb 26, 2026 · 1h 3m · mad

AI That Can Prove It’s Right: Verification as the Missing Layer in AI — Carina Hong

Corinna Hong · 50m spoken Matt Turck · 9m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of The MAD Podcast, host Matt Turck interviews Corinna Hong, Founder and CEO of Axiom Math, about how formal automated theorem proving in Lean eliminates AI hallucinations to create provably accurate reasoning systems. They discuss Axiom's milestones in acing competitive math exams and solving open research conjectures, as well as the company's broader roadmap for formally verified software engineering and scientific discovery.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 15.9% of the talking time here. How this is scored →

Matt as informed peer 2.6 Guest teaching 4.0 Guest disagreement 0.8 Matt pushing back 0.7
05100:0015:0030:0045:001:00:001:23–4:05 · Matt as informed peer 1/10 The Need for an AI Mathematician The host asks basic introductory questions setting up Axiom's mission and asking for definitions of math competition levels like the Putnam. The guest explains the two axes of mathematical difficulty (creativity vs abstraction) cleanly without friction.4:05–6:59 · Matt as informed peer 2/10 Autonomously Solving Open Research Conjectures The host presents an uneducated framing regarding top mathematicians relying on intuition and serendipity. The guest gently nuances this by explaining how standard routines vs eureka moments function in research math.6:59–8:59 · Matt as informed peer 4/10 "Move 37" in Math: Lean Proofs vs. Human Solutions The host brings up the famous Move 37 concept from AlphaGo to ask if AI solves math in non-human ways. The guest explains how Lean proofs lean heavily into routine bookkeeping and casework that human mathematicians shy away from.8:59–14:27 · Matt as informed peer 3/10 Understanding Lean: Code as Mathematical Proof The host prompts for a definition of Lean and asks how Axiom compares to OpenAI and DeepMind's IMO systems. The guest educates the host on Curry-Howard correspondence, the history of automated theorem proving, and AlphaGeometry.14:27–17:37 · Matt as informed peer 4/10 Informal Natural Language vs. Formal Machine Code in Math The host attempts to categorize DeepMind as informal brute-force compared to Axiom's neuro-symbolic approach. The guest corrects the host's premise, pointing out that AlphaProof was in fact a formal theorem proving system.17:37–20:18 · Matt as informed peer 3/10 Verification Beyond Final Answers: RLVR in Proof-Based Math The host asks about reinforcement learning with verifiable rewards (RLVR). The guest explains the flaw of numerical answer verification (reward hacking) versus step-by-step proof verification in adult research math.20:18–23:22 · Matt as informed peer 3/10 Eliminating Hallucinations: Expanding Verification to Code and Beyond The host asks whether each domain requires its own Lean equivalent as Axiom expands beyond math. The guest details why strongly typed languages like Rust are structurally much closer to Lean than natural language.23:22–29:30 · Matt as informed peer 3/10 The Future Roadmap: From Math Verification to AI for Science The host asks where the threshold lies between scientific code and general text before pivoting to the guest's early life in China. The guest draws an analogy between her experience at Ross Math Camp and Axiom Prover's skill library.29:39–33:52 · Matt as informed peer 2/10 Academic Journey: MIT, Rhodes Scholarship, and Neuroscience at UCL Gatsby The host asks about the guest's unique background combining MIT, a Rhodes scholarship in neuroscience at UCL Gatsby, and a JD/PhD at Stanford. The guest explains her interdisciplinary path collaboratively.33:52–40:50 · Matt as informed peer 3/10 Axiom Architecture: Conjecture, Prover, and Knowledge Base The host asks about product architecture, LLM usage, synthetic data, and compute requirements. The guest uses a ocean sailing metaphor to explain conjecture, prover, and knowledge base before detailing formal synthetic data fuzzing.40:50–42:58 · Matt as informed peer 3/10 Product Maturity, Research Frontiers, and Expert Hiring The host inquires about current product maturity and the split between research and engineering. The guest explains how the team pushes research frontiers while selectively hardening capabilities.42:58–45:11 · Matt as informed peer 3/10 Scaling Reasoning Complexity: From Putnam Problems to Frontier Math The host asks if there is an Everest in math for AI to tackle. The guest educates the host on journal publication tiers and details the jump in reasoning tree complexity from 40 nodes to thousands of nodes.45:11–47:25 · Matt as informed peer 4/10 Scaling Up vs. Scaling Out Across Mathematical and Physical Domains The host asks if AI can win a Fields Medal. The guest reframes the question, stating that mathematicians aim for the shortlist rather than the medal itself due to human factors, prompting the host to press on why.47:25–49:29 · Matt as informed peer 2/10 Catalyzing a Mathematical Renaissance and Scientific Breakthroughs The host asks about AI generating groundbreaking scientific discoveries. The guest outlines shifting from a scarcity mindset around elite mathematical human labor to an era of AI-driven mathematical abundance.49:29–53:00 · Matt as informed peer 2/10 Theoretical Physics, 'Last Mile' Precision, and Discovery Cycles The host prompts on timelines to reach this world. The guest discusses the value of last-mile precision and contrasts her view that math is code with CTO Shubo's view that code is math.53:00–55:47 · Matt as informed peer 1/10 Code as Math: High-Stakes Systems and Formal Verification The host listens as the guest explains high-stakes industrial applications like GPU chip design and aerospace where partial verification is useless and formal guarantees are required.55:47–59:42 · Matt as informed peer 3/10 Mathematical Skepticism vs. Acceptance of Formal Verifiers The host asks whether the broader math community fears replacement by AI. The guest explains that mathematicians reject uncheckable natural language LLM proofs but welcome Lean formal certificates.59:46–1:03:22 · Matt as informed peer 3/10 Evaluating Talent and Hiring for Taste The host asks intelligent questions about evaluating talent, hiring for taste, and assembling world-class founding mathematicians. The guest shares her whale-frequency metaphor for recruiting top talent.1:03:22–1:03:59 · Matt as informed peer 0/10 Concluding the Interview and Expressing Gratitude Standard conversational sign-off and podcast outro; no technical exchange or pushback.1:23–4:05 · Guest teaching 3/10 The Need for an AI Mathematician The host asks basic introductory questions setting up Axiom's mission and asking for definitions of math competition levels like the Putnam. The guest explains the two axes of mathematical difficulty (creativity vs abstraction) cleanly without friction.4:05–6:59 · Guest teaching 4/10 Autonomously Solving Open Research Conjectures The host presents an uneducated framing regarding top mathematicians relying on intuition and serendipity. The guest gently nuances this by explaining how standard routines vs eureka moments function in research math.6:59–8:59 · Guest teaching 4/10 "Move 37" in Math: Lean Proofs vs. Human Solutions The host brings up the famous Move 37 concept from AlphaGo to ask if AI solves math in non-human ways. The guest explains how Lean proofs lean heavily into routine bookkeeping and casework that human mathematicians shy away from.8:59–14:27 · Guest teaching 6/10 Understanding Lean: Code as Mathematical Proof The host prompts for a definition of Lean and asks how Axiom compares to OpenAI and DeepMind's IMO systems. The guest educates the host on Curry-Howard correspondence, the history of automated theorem proving, and AlphaGeometry.14:27–17:37 · Guest teaching 5/10 Informal Natural Language vs. Formal Machine Code in Math The host attempts to categorize DeepMind as informal brute-force compared to Axiom's neuro-symbolic approach. The guest corrects the host's premise, pointing out that AlphaProof was in fact a formal theorem proving system.17:37–20:18 · Guest teaching 5/10 Verification Beyond Final Answers: RLVR in Proof-Based Math The host asks about reinforcement learning with verifiable rewards (RLVR). The guest explains the flaw of numerical answer verification (reward hacking) versus step-by-step proof verification in adult research math.20:18–23:22 · Guest teaching 4/10 Eliminating Hallucinations: Expanding Verification to Code and Beyond The host asks whether each domain requires its own Lean equivalent as Axiom expands beyond math. The guest details why strongly typed languages like Rust are structurally much closer to Lean than natural language.23:22–29:30 · Guest teaching 4/10 The Future Roadmap: From Math Verification to AI for Science The host asks where the threshold lies between scientific code and general text before pivoting to the guest's early life in China. The guest draws an analogy between her experience at Ross Math Camp and Axiom Prover's skill library.29:39–33:52 · Guest teaching 3/10 Academic Journey: MIT, Rhodes Scholarship, and Neuroscience at UCL Gatsby The host asks about the guest's unique background combining MIT, a Rhodes scholarship in neuroscience at UCL Gatsby, and a JD/PhD at Stanford. The guest explains her interdisciplinary path collaboratively.33:52–40:50 · Guest teaching 5/10 Axiom Architecture: Conjecture, Prover, and Knowledge Base The host asks about product architecture, LLM usage, synthetic data, and compute requirements. The guest uses a ocean sailing metaphor to explain conjecture, prover, and knowledge base before detailing formal synthetic data fuzzing.40:50–42:58 · Guest teaching 3/10 Product Maturity, Research Frontiers, and Expert Hiring The host inquires about current product maturity and the split between research and engineering. The guest explains how the team pushes research frontiers while selectively hardening capabilities.42:58–45:11 · Guest teaching 5/10 Scaling Reasoning Complexity: From Putnam Problems to Frontier Math The host asks if there is an Everest in math for AI to tackle. The guest educates the host on journal publication tiers and details the jump in reasoning tree complexity from 40 nodes to thousands of nodes.45:11–47:25 · Guest teaching 5/10 Scaling Up vs. Scaling Out Across Mathematical and Physical Domains The host asks if AI can win a Fields Medal. The guest reframes the question, stating that mathematicians aim for the shortlist rather than the medal itself due to human factors, prompting the host to press on why.47:25–49:29 · Guest teaching 4/10 Catalyzing a Mathematical Renaissance and Scientific Breakthroughs The host asks about AI generating groundbreaking scientific discoveries. The guest outlines shifting from a scarcity mindset around elite mathematical human labor to an era of AI-driven mathematical abundance.49:29–53:00 · Guest teaching 4/10 Theoretical Physics, 'Last Mile' Precision, and Discovery Cycles The host prompts on timelines to reach this world. The guest discusses the value of last-mile precision and contrasts her view that math is code with CTO Shubo's view that code is math.53:00–55:47 · Guest teaching 4/10 Code as Math: High-Stakes Systems and Formal Verification The host listens as the guest explains high-stakes industrial applications like GPU chip design and aerospace where partial verification is useless and formal guarantees are required.55:47–59:42 · Guest teaching 4/10 Mathematical Skepticism vs. Acceptance of Formal Verifiers The host asks whether the broader math community fears replacement by AI. The guest explains that mathematicians reject uncheckable natural language LLM proofs but welcome Lean formal certificates.59:46–1:03:22 · Guest teaching 3/10 Evaluating Talent and Hiring for Taste The host asks intelligent questions about evaluating talent, hiring for taste, and assembling world-class founding mathematicians. The guest shares her whale-frequency metaphor for recruiting top talent.1:03:22–1:03:59 · Guest teaching 0/10 Concluding the Interview and Expressing Gratitude Standard conversational sign-off and podcast outro; no technical exchange or pushback.1:23–4:05 · Guest disagreement 0/10 The Need for an AI Mathematician The host asks basic introductory questions setting up Axiom's mission and asking for definitions of math competition levels like the Putnam. The guest explains the two axes of mathematical difficulty (creativity vs abstraction) cleanly without friction.4:05–6:59 · Guest disagreement 1/10 Autonomously Solving Open Research Conjectures The host presents an uneducated framing regarding top mathematicians relying on intuition and serendipity. The guest gently nuances this by explaining how standard routines vs eureka moments function in research math.6:59–8:59 · Guest disagreement 1/10 "Move 37" in Math: Lean Proofs vs. Human Solutions The host brings up the famous Move 37 concept from AlphaGo to ask if AI solves math in non-human ways. The guest explains how Lean proofs lean heavily into routine bookkeeping and casework that human mathematicians shy away from.8:59–14:27 · Guest disagreement 1/10 Understanding Lean: Code as Mathematical Proof The host prompts for a definition of Lean and asks how Axiom compares to OpenAI and DeepMind's IMO systems. The guest educates the host on Curry-Howard correspondence, the history of automated theorem proving, and AlphaGeometry.14:27–17:37 · Guest disagreement 2/10 Informal Natural Language vs. Formal Machine Code in Math The host attempts to categorize DeepMind as informal brute-force compared to Axiom's neuro-symbolic approach. The guest corrects the host's premise, pointing out that AlphaProof was in fact a formal theorem proving system.17:37–20:18 · Guest disagreement 1/10 Verification Beyond Final Answers: RLVR in Proof-Based Math The host asks about reinforcement learning with verifiable rewards (RLVR). The guest explains the flaw of numerical answer verification (reward hacking) versus step-by-step proof verification in adult research math.20:18–23:22 · Guest disagreement 1/10 Eliminating Hallucinations: Expanding Verification to Code and Beyond The host asks whether each domain requires its own Lean equivalent as Axiom expands beyond math. The guest details why strongly typed languages like Rust are structurally much closer to Lean than natural language.23:22–29:30 · Guest disagreement 1/10 The Future Roadmap: From Math Verification to AI for Science The host asks where the threshold lies between scientific code and general text before pivoting to the guest's early life in China. The guest draws an analogy between her experience at Ross Math Camp and Axiom Prover's skill library.29:39–33:52 · Guest disagreement 0/10 Academic Journey: MIT, Rhodes Scholarship, and Neuroscience at UCL Gatsby The host asks about the guest's unique background combining MIT, a Rhodes scholarship in neuroscience at UCL Gatsby, and a JD/PhD at Stanford. The guest explains her interdisciplinary path collaboratively.33:52–40:50 · Guest disagreement 1/10 Axiom Architecture: Conjecture, Prover, and Knowledge Base The host asks about product architecture, LLM usage, synthetic data, and compute requirements. The guest uses a ocean sailing metaphor to explain conjecture, prover, and knowledge base before detailing formal synthetic data fuzzing.40:50–42:58 · Guest disagreement 0/10 Product Maturity, Research Frontiers, and Expert Hiring The host inquires about current product maturity and the split between research and engineering. The guest explains how the team pushes research frontiers while selectively hardening capabilities.42:58–45:11 · Guest disagreement 1/10 Scaling Reasoning Complexity: From Putnam Problems to Frontier Math The host asks if there is an Everest in math for AI to tackle. The guest educates the host on journal publication tiers and details the jump in reasoning tree complexity from 40 nodes to thousands of nodes.45:11–47:25 · Guest disagreement 2/10 Scaling Up vs. Scaling Out Across Mathematical and Physical Domains The host asks if AI can win a Fields Medal. The guest reframes the question, stating that mathematicians aim for the shortlist rather than the medal itself due to human factors, prompting the host to press on why.47:25–49:29 · Guest disagreement 1/10 Catalyzing a Mathematical Renaissance and Scientific Breakthroughs The host asks about AI generating groundbreaking scientific discoveries. The guest outlines shifting from a scarcity mindset around elite mathematical human labor to an era of AI-driven mathematical abundance.49:29–53:00 · Guest disagreement 1/10 Theoretical Physics, 'Last Mile' Precision, and Discovery Cycles The host prompts on timelines to reach this world. The guest discusses the value of last-mile precision and contrasts her view that math is code with CTO Shubo's view that code is math.53:00–55:47 · Guest disagreement 0/10 Code as Math: High-Stakes Systems and Formal Verification The host listens as the guest explains high-stakes industrial applications like GPU chip design and aerospace where partial verification is useless and formal guarantees are required.55:47–59:42 · Guest disagreement 1/10 Mathematical Skepticism vs. Acceptance of Formal Verifiers The host asks whether the broader math community fears replacement by AI. The guest explains that mathematicians reject uncheckable natural language LLM proofs but welcome Lean formal certificates.59:46–1:03:22 · Guest disagreement 0/10 Evaluating Talent and Hiring for Taste The host asks intelligent questions about evaluating talent, hiring for taste, and assembling world-class founding mathematicians. The guest shares her whale-frequency metaphor for recruiting top talent.1:03:22–1:03:59 · Guest disagreement 0/10 Concluding the Interview and Expressing Gratitude Standard conversational sign-off and podcast outro; no technical exchange or pushback.1:23–4:05 · Matt pushing back 0/10 The Need for an AI Mathematician The host asks basic introductory questions setting up Axiom's mission and asking for definitions of math competition levels like the Putnam. The guest explains the two axes of mathematical difficulty (creativity vs abstraction) cleanly without friction.4:05–6:59 · Matt pushing back 1/10 Autonomously Solving Open Research Conjectures The host presents an uneducated framing regarding top mathematicians relying on intuition and serendipity. The guest gently nuances this by explaining how standard routines vs eureka moments function in research math.6:59–8:59 · Matt pushing back 1/10 "Move 37" in Math: Lean Proofs vs. Human Solutions The host brings up the famous Move 37 concept from AlphaGo to ask if AI solves math in non-human ways. The guest explains how Lean proofs lean heavily into routine bookkeeping and casework that human mathematicians shy away from.8:59–14:27 · Matt pushing back 1/10 Understanding Lean: Code as Mathematical Proof The host prompts for a definition of Lean and asks how Axiom compares to OpenAI and DeepMind's IMO systems. The guest educates the host on Curry-Howard correspondence, the history of automated theorem proving, and AlphaGeometry.14:27–17:37 · Matt pushing back 2/10 Informal Natural Language vs. Formal Machine Code in Math The host attempts to categorize DeepMind as informal brute-force compared to Axiom's neuro-symbolic approach. The guest corrects the host's premise, pointing out that AlphaProof was in fact a formal theorem proving system.17:37–20:18 · Matt pushing back 1/10 Verification Beyond Final Answers: RLVR in Proof-Based Math The host asks about reinforcement learning with verifiable rewards (RLVR). The guest explains the flaw of numerical answer verification (reward hacking) versus step-by-step proof verification in adult research math.20:18–23:22 · Matt pushing back 1/10 Eliminating Hallucinations: Expanding Verification to Code and Beyond The host asks whether each domain requires its own Lean equivalent as Axiom expands beyond math. The guest details why strongly typed languages like Rust are structurally much closer to Lean than natural language.23:22–29:30 · Matt pushing back 1/10 The Future Roadmap: From Math Verification to AI for Science The host asks where the threshold lies between scientific code and general text before pivoting to the guest's early life in China. The guest draws an analogy between her experience at Ross Math Camp and Axiom Prover's skill library.29:39–33:52 · Matt pushing back 0/10 Academic Journey: MIT, Rhodes Scholarship, and Neuroscience at UCL Gatsby The host asks about the guest's unique background combining MIT, a Rhodes scholarship in neuroscience at UCL Gatsby, and a JD/PhD at Stanford. The guest explains her interdisciplinary path collaboratively.33:52–40:50 · Matt pushing back 1/10 Axiom Architecture: Conjecture, Prover, and Knowledge Base The host asks about product architecture, LLM usage, synthetic data, and compute requirements. The guest uses a ocean sailing metaphor to explain conjecture, prover, and knowledge base before detailing formal synthetic data fuzzing.40:50–42:58 · Matt pushing back 0/10 Product Maturity, Research Frontiers, and Expert Hiring The host inquires about current product maturity and the split between research and engineering. The guest explains how the team pushes research frontiers while selectively hardening capabilities.42:58–45:11 · Matt pushing back 1/10 Scaling Reasoning Complexity: From Putnam Problems to Frontier Math The host asks if there is an Everest in math for AI to tackle. The guest educates the host on journal publication tiers and details the jump in reasoning tree complexity from 40 nodes to thousands of nodes.45:11–47:25 · Matt pushing back 2/10 Scaling Up vs. Scaling Out Across Mathematical and Physical Domains The host asks if AI can win a Fields Medal. The guest reframes the question, stating that mathematicians aim for the shortlist rather than the medal itself due to human factors, prompting the host to press on why.47:25–49:29 · Matt pushing back 0/10 Catalyzing a Mathematical Renaissance and Scientific Breakthroughs The host asks about AI generating groundbreaking scientific discoveries. The guest outlines shifting from a scarcity mindset around elite mathematical human labor to an era of AI-driven mathematical abundance.49:29–53:00 · Matt pushing back 1/10 Theoretical Physics, 'Last Mile' Precision, and Discovery Cycles The host prompts on timelines to reach this world. The guest discusses the value of last-mile precision and contrasts her view that math is code with CTO Shubo's view that code is math.53:00–55:47 · Matt pushing back 0/10 Code as Math: High-Stakes Systems and Formal Verification The host listens as the guest explains high-stakes industrial applications like GPU chip design and aerospace where partial verification is useless and formal guarantees are required.55:47–59:42 · Matt pushing back 1/10 Mathematical Skepticism vs. Acceptance of Formal Verifiers The host asks whether the broader math community fears replacement by AI. The guest explains that mathematicians reject uncheckable natural language LLM proofs but welcome Lean formal certificates.59:46–1:03:22 · Matt pushing back 0/10 Evaluating Talent and Hiring for Taste The host asks intelligent questions about evaluating talent, hiring for taste, and assembling world-class founding mathematicians. The guest shares her whale-frequency metaphor for recruiting top talent.1:03:22–1:03:59 · Matt pushing back 0/10 Concluding the Interview and Expressing Gratitude Standard conversational sign-off and podcast outro; no technical exchange or pushback.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 39% · guest 61%0:00 · Matt 39% · guest 61%3:00 · Matt 23.8% · guest 76.2%3:00 · Matt 23.8% · guest 76.2%6:00 · Matt 12.2% · guest 87.8%6:00 · Matt 12.2% · guest 87.8%9:00 · Matt 15.2% · guest 84.8%9:00 · Matt 15.2% · guest 84.8%12:00 · Matt 1.1% · guest 98.9%12:00 · Matt 1.1% · guest 98.9%15:00 · Matt 23.5% · guest 76.5%15:00 · Matt 23.5% · guest 76.5%18:00 · Matt 14.5% · guest 85.5%18:00 · Matt 14.5% · guest 85.5%21:00 · Matt 14.7% · guest 85.3%21:00 · Matt 14.7% · guest 85.3%24:00 · Matt 21.3% · guest 78.7%24:00 · Matt 21.3% · guest 78.7%27:00 · Matt 12.8% · guest 87.2%27:00 · Matt 12.8% · guest 87.2%30:00 · Matt 16.1% · guest 83.9%30:00 · Matt 16.1% · guest 83.9%33:00 · Matt 17.7% · guest 82.3%33:00 · Matt 17.7% · guest 82.3%36:00 · Matt 17.2% · guest 82.8%36:00 · Matt 17.2% · guest 82.8%39:00 · Matt 15.9% · guest 84.1%39:00 · Matt 15.9% · guest 84.1%42:00 · Matt 8.3% · guest 91.7%42:00 · Matt 8.3% · guest 91.7%45:00 · Matt 27.5% · guest 72.5%45:00 · Matt 27.5% · guest 72.5%48:00 · Matt 0% · guest 100%48:00 · Matt 0% · guest 100%51:00 · Matt 3.1% · guest 96.9%51:00 · Matt 3.1% · guest 96.9%54:00 · Matt 8.1% · guest 91.9%54:00 · Matt 8.1% · guest 91.9%57:00 · Matt 18.3% · guest 81.7%57:00 · Matt 18.3% · guest 81.7%1:00:00 · Matt 13% · guest 87%1:00:00 · Matt 13% · guest 87%1:03:00 · Matt 53.1% · guest 46.9%1:03:00 · Matt 53.1% · guest 46.9%
Sharpest disagreement ▶ 46:36 Guest rejects Fields Medal premise

When the host asks if AI can win a Fields Medal, the guest explicitly reframes the benchmark, asserting that real mathematicians celebrate reaching the shortlist rather than winning the medal itself.

Hardest push from Matt ▶ 47:07 Host presses on Fields Medal shortlist reframe

The host refuses to let the guest's reframe pass without interjecting to ask why shortlist status is the real benchmark and inquiring if political factors enter the decision.

Biggest teaching moment ▶ 16:07 Guest corrects host on Google DeepMind's formal system

When the host attempts to frame Google DeepMind's IMO work as informal brute-force relative to Axiom, the guest directly corrects him by clarifying that AlphaProof was a formal theorem proving system.

Matt holds his own ▶ 6:58 Host introduces DeepMind Move 37 analogy

The host demonstrates strong domain knowledge in AI history by drawing a comparison to AlphaGo's famous Move 37 to ask whether AI math proofs exhibit alien reasoning patterns.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
The Need for an AI Mathematician 1300 The host asks basic introductory questions setting up Axiom's mission and asking for definitions of math competition levels like the Putnam. The guest explains the two axes of mathematical difficulty (creativity vs abstraction) cleanly without friction.
Autonomously Solving Open Research Conjectures 2411 The host presents an uneducated framing regarding top mathematicians relying on intuition and serendipity. The guest gently nuances this by explaining how standard routines vs eureka moments function in research math.
"Move 37" in Math: Lean Proofs vs. Human Solutions 4411 The host brings up the famous Move 37 concept from AlphaGo to ask if AI solves math in non-human ways. The guest explains how Lean proofs lean heavily into routine bookkeeping and casework that human mathematicians shy away from.
Understanding Lean: Code as Mathematical Proof 3611 The host prompts for a definition of Lean and asks how Axiom compares to OpenAI and DeepMind's IMO systems. The guest educates the host on Curry-Howard correspondence, the history of automated theorem proving, and AlphaGeometry.
Informal Natural Language vs. Formal Machine Code in Math 4522 The host attempts to categorize DeepMind as informal brute-force compared to Axiom's neuro-symbolic approach. The guest corrects the host's premise, pointing out that AlphaProof was in fact a formal theorem proving system.
Verification Beyond Final Answers: RLVR in Proof-Based Math 3511 The host asks about reinforcement learning with verifiable rewards (RLVR). The guest explains the flaw of numerical answer verification (reward hacking) versus step-by-step proof verification in adult research math.
Eliminating Hallucinations: Expanding Verification to Code and Beyond 3411 The host asks whether each domain requires its own Lean equivalent as Axiom expands beyond math. The guest details why strongly typed languages like Rust are structurally much closer to Lean than natural language.
The Future Roadmap: From Math Verification to AI for Science 3411 The host asks where the threshold lies between scientific code and general text before pivoting to the guest's early life in China. The guest draws an analogy between her experience at Ross Math Camp and Axiom Prover's skill library.
Academic Journey: MIT, Rhodes Scholarship, and Neuroscience at UCL Gatsby 2300 The host asks about the guest's unique background combining MIT, a Rhodes scholarship in neuroscience at UCL Gatsby, and a JD/PhD at Stanford. The guest explains her interdisciplinary path collaboratively.
Axiom Architecture: Conjecture, Prover, and Knowledge Base 3511 The host asks about product architecture, LLM usage, synthetic data, and compute requirements. The guest uses a ocean sailing metaphor to explain conjecture, prover, and knowledge base before detailing formal synthetic data fuzzing.
Product Maturity, Research Frontiers, and Expert Hiring 3300 The host inquires about current product maturity and the split between research and engineering. The guest explains how the team pushes research frontiers while selectively hardening capabilities.
Scaling Reasoning Complexity: From Putnam Problems to Frontier Math 3511 The host asks if there is an Everest in math for AI to tackle. The guest educates the host on journal publication tiers and details the jump in reasoning tree complexity from 40 nodes to thousands of nodes.
Scaling Up vs. Scaling Out Across Mathematical and Physical Domains 4522 The host asks if AI can win a Fields Medal. The guest reframes the question, stating that mathematicians aim for the shortlist rather than the medal itself due to human factors, prompting the host to press on why.
Catalyzing a Mathematical Renaissance and Scientific Breakthroughs 2410 The host asks about AI generating groundbreaking scientific discoveries. The guest outlines shifting from a scarcity mindset around elite mathematical human labor to an era of AI-driven mathematical abundance.
Theoretical Physics, 'Last Mile' Precision, and Discovery Cycles 2411 The host prompts on timelines to reach this world. The guest discusses the value of last-mile precision and contrasts her view that math is code with CTO Shubo's view that code is math.
Code as Math: High-Stakes Systems and Formal Verification 1400 The host listens as the guest explains high-stakes industrial applications like GPU chip design and aerospace where partial verification is useless and formal guarantees are required.
Mathematical Skepticism vs. Acceptance of Formal Verifiers 3411 The host asks whether the broader math community fears replacement by AI. The guest explains that mathematicians reject uncheckable natural language LLM proofs but welcome Lean formal certificates.
Evaluating Talent and Hiring for Taste 3300 The host asks intelligent questions about evaluating talent, hiring for taste, and assembling world-class founding mathematicians. The guest shares her whale-frequency metaphor for recruiting top talent.
Concluding the Interview and Expressing Gratitude 0000 Standard conversational sign-off and podcast outro; no technical exchange or pushback.

Statements from this episode (31)

Insight
Mathematical AI unlocks solutions for verification and optimization
“I think through solving math, we also realize that it can solve a lot of other problems, such as verification, such as optimization, et cetera.”
Corinna Hong Feb 26, 2026 ▶ 1:45
Insight
Mathematical difficulty spans solution creativity and object abstraction
“There are actually two axis of difficulty. One is how creative the solution is, and The other one, roughly speaking, is how abstract the mathematical object is.”
Corinna Hong Feb 26, 2026 ▶ 2:23
Assertion Supported
Only five humans have achieved a perfect Putnam score in 100 years
“Over the 100 year history of Putnam, there's only five human perfect scores.”
Corinna Hong Feb 26, 2026 ▶ 3:25
Assertion Supported
AxiomProver achieved a perfect score on the 2025 Putnam math exam
“Eight within the time limit, and then 12 out of 12.”
Corinna Hong Feb 26, 2026 ▶ 4:02
Assertion Supported
AxiomProver has solved four research-level open mathematical problems
“Recently, I think like a couple of weeks ago, we just announced that action prover solved these four research level open problems.”
Corinna Hong Feb 26, 2026 ▶ 4:50
Assertion Not checkable as stated
AxiomProver is the first AI to solve research conjectures end-to-end
“It's probably the first AI to solve a research conjecture completely end-to-end and self-verifies. That means the output are fully verified, a hundred percent correct.”
Corinna Hong Feb 26, 2026 ▶ 4:56
Prediction Not checkable as stated
Today's AI can solve math problems that take human researchers months
“I think that we are at a threshold of mathematical renaissance, which is to realize that there are so many unsolved problems that will currently take, say, researchers months to crack, or even technical lemmas in those really longstanding conjectures that we b…”
Corinna Hong Feb 26, 2026 ▶ 6:14
Assertion Not checkable as stated
Axiom Math's AI system proves multiple open research conjectures every week
“We actually have actually a few more research conjectures that's being proven every week just by the supply of mathematicians you know, from the world, and we try to put those problems into use.”
Corinna Hong Feb 26, 2026 ▶ 6:47
Assertion Not checkable as stated
Axiom's AI generated Putnam math solutions that diverge from human proofs
“So we actually analyzed all 12 problem solutions of the Putnam exam, and we found that a lot of the solutions actually differ from the human solution.”
Corinna Hong Feb 26, 2026 ▶ 8:22
Insight
Lean-based AI provers favor mechanistic arguments over clever human-style solutions
“Because it is a, you know, lean based system, it is really good at sort of routine bookkeeping, and it will actually choose a lot of the more mechanistic, you know, arguments over the ones that require like a clever, say one picture solution.”
Corinna Hong Feb 26, 2026 ▶ 8:30
Assertion Not checkable as stated
Auto-formalization is harder than translating between two programming languages
“And auto formalization, which is the sort of capability of converting the natural language reasoning to say the formal language. And that's harder than translation because it's different than say translating between two programming languages. You're translatin…”
Corinna Hong Feb 26, 2026 ▶ 15:29
Assertion Not checkable as stated
Translating formal Lean code into English is easier than auto-formalization
“There's also auto-informalization, which is kind of translate back, I mean, from Lean to English. That's easier than auto-formalization, because most of the machines' AI have seen a lot more English than Lean.”
Corinna Hong Feb 26, 2026 ▶ 15:55
Disclosure
Axiom Math focuses on post-training reinforcement learning to achieve performance gains
“And I think that we shouldn't do pre-training. We shouldn't try to just only train from scratch. I think we're kind of focusing on post-training reinforcement learning can potentially get us better performance gain.”
Corinna Hong Feb 26, 2026 ▶ 17:26
Assertion Not checkable as stated
Numerical AI benchmarks fail to evaluate underlying logical reasoning capabilities
“Like, you know, we have seen from, say, Frontier Math and other benchmark, which only compels a numerical answer that it doesn't actually necessarily reflect the model's capability in the logical reasoning.”
Corinna Hong Feb 26, 2026 ▶ 18:17
Insight
True AI reasoning engines require verifiable rewards for intermediate proof steps
“If you want to have a reasoning engine that really truly masters at logic and mathematical reasoning, then you need to somehow get verifiable reward for the proof steps.”
Corinna Hong Feb 26, 2026 ▶ 19:52
Insight
Mathematical reasoning is the foundational reasoning layer for AGI
“Our worldview is math reasoning is a true reasoning layer of AGI.”
Corinna Hong Feb 26, 2026 ▶ 21:26
Insight
The Lean theorem prover language is closer to Rust than to English
“I think the sort of gap between, say, for example, Lean and another, like, strongly typed language like Rust is a lot closer than the gap between Lean and English.”
Corinna Hong Feb 26, 2026 ▶ 22:58
Assertion Contradicted
The open-source Lean dataset contains only tens of millions of tokens
“It's only two-digit million number of tokens out there in the open, open world.”
Corinna Hong Feb 26, 2026 ▶ 28:35
Disclosure
AxiomProver self-improves by adding generated mathematical proofs to its skill library
“Action Prover learns to prove things, and it kind of self-improved in a way where all the things that it proved got fed back into it, into a kind of a skill library.”
Corinna Hong Feb 26, 2026 ▶ 28:40
Disclosure
Axiom Math's architecture combines a conjecture generator, prover, and knowledge base
“Our kind of very broad vision is that we are going to have a conjecture. We're going to have a prover. And then there is knowledge base.”
Corinna Hong Feb 26, 2026 ▶ 34:07
Disclosure
Axiom Math will release its Lean tools on a public API
“We are actually gonna release them on a public, like, API on these, all these dozen of pools. Very, very soon. Beginning of March.”
Corinna Hong Feb 26, 2026 ▶ 36:40
Assertion Open · timeframe Feb 2029
Axiom's proof verifier is 100 times faster than open-source alternatives
“So a lot of the sort of like verify, verify proof is actually, you know, one of our prover tools that's about to be released, and that's actually a hundred times faster than The other counterparts that are the open source, like effort, cloud comparator.”
Corinna Hong Feb 26, 2026 ▶ 37:32
Assertion Not checkable as stated
Axiom Math tested transfer learning from mathematical reasoning to code verification
“So from Putnam Perfect Score, that was four months in, then two months later was the four research conjectures, and then, you know, during this middle, we also have tested something that is transfer learning from math to code verification, so another evaluatio…”
Corinna Hong Feb 26, 2026 ▶ 41:44
Assertion Not checkable as stated
AxiomProver autonomously proves theorems publishable in major mathematical journals
“Currently the batch of papers, Axiom Prover has autonomously proven and mathematicians have written You can probably get into Journal of Number Theory, Journal of Algebra, like that level.”
Corinna Hong Feb 26, 2026 ▶ 43:12
Assertion Not checkable as stated
Hard internal research problems require reasoning trees with thousands of nodes
“So on the easy end of the pun-end problem, we have 40 nodes. On the hard end of research questions in-house, we currently have a research problem with thousands of nodes.”
Corinna Hong Feb 26, 2026 ▶ 44:05
Disclosure
Axiom Math aims to solve a Fields Medal shortlist-worthy problem using AI
“I think that we really want Accent Prover to be able to solve one long-standing problem in mathematics that you can objectively, objectively say, even though if it's an AI, you know, or double-blind, whatever, that will be in the shortlist.”
Corinna Hong Feb 26, 2026 ▶ 46:49
Prediction Not checkable as stated
AxiomProver could eventually solve the majority of human mathematical conjectures
“Everything that human mind Can conjecture, find interesting, find tasteful, could be solved by, hopefully, majority of them by accent prover.”
Corinna Hong Feb 26, 2026 ▶ 48:56
Assertion Not checkable as stated
Domain experts grade the 2025 Putnam Competition as harder than IMO 2025
“Pundum 20 25 is by a lot of sort of experts grading harder than IMO 20 25, so it's the hardest reward Maths Olympiae test”
Corinna Hong Feb 26, 2026 ▶ 54:30
Insight
Math community AI skepticism stems from unverifiable informal language model outputs
“A lot of the, like, adverse reaction about from the math community about AI is actually coming from the fact that they cannot verify an informal solution.”
Corinna Hong Feb 26, 2026 ▶ 56:01
Assertion Not checkable as stated
Almost all younger-generation mathematicians accept Lean as a formal proof verifier
“Almost all of the new school of mathematicians are accepting Lean.”
Corinna Hong Feb 26, 2026 ▶ 56:42
Insight
Generation and verification loops are the next major frontier of AI
“We still feel like we cannot fully elaborate and emphasize the thing that we are seeing that is the next frontier of AI. That is a generation and verification loop. That is the discovery of verified knowledge.”
Corinna Hong Feb 26, 2026 ▶ 1:03:09
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.