Feb 26, 2026 · 1h 3m · mad
AI That Can Prove It’s Right: Verification as the Missing Layer in AI — Carina Hong
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of The MAD Podcast, host Matt Turck interviews Corinna Hong, Founder and CEO of Axiom Math, about how formal automated theorem proving in Lean eliminates AI hallucinations to create provably accurate reasoning systems. They discuss Axiom's milestones in acing competitive math exams and solving open research conjectures, as well as the company's broader roadmap for formally verified software engineering and scientific discovery.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 15.9% of the talking time here. How this is scored →
speaking balance: gold is Matt, purple is the guest (3 minute bins)
When the host asks if AI can win a Fields Medal, the guest explicitly reframes the benchmark, asserting that real mathematicians celebrate reaching the shortlist rather than winning the medal itself.
Hardest push from Matt ▶ 47:07 Host presses on Fields Medal shortlist reframeThe host refuses to let the guest's reframe pass without interjecting to ask why shortlist status is the real benchmark and inquiring if political factors enter the decision.
Biggest teaching moment ▶ 16:07 Guest corrects host on Google DeepMind's formal systemWhen the host attempts to frame Google DeepMind's IMO work as informal brute-force relative to Axiom, the guest directly corrects him by clarifying that AlphaProof was a formal theorem proving system.
Matt holds his own ▶ 6:58 Host introduces DeepMind Move 37 analogyThe host demonstrates strong domain knowledge in AI history by drawing a comparison to AlphaGo's famous Move 37 to ask whether AI math proofs exhibit alien reasoning patterns.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Matt as informed peer | Guest teaching | Guest disagreement | Matt pushing back | Why |
|---|---|---|---|---|---|---|
| The Need for an AI Mathematician | 1 | 3 | 0 | 0 | The host asks basic introductory questions setting up Axiom's mission and asking for definitions of math competition levels like the Putnam. The guest explains the two axes of mathematical difficulty (creativity vs abstraction) cleanly without friction. | |
| Autonomously Solving Open Research Conjectures | 2 | 4 | 1 | 1 | The host presents an uneducated framing regarding top mathematicians relying on intuition and serendipity. The guest gently nuances this by explaining how standard routines vs eureka moments function in research math. | |
| "Move 37" in Math: Lean Proofs vs. Human Solutions | 4 | 4 | 1 | 1 | The host brings up the famous Move 37 concept from AlphaGo to ask if AI solves math in non-human ways. The guest explains how Lean proofs lean heavily into routine bookkeeping and casework that human mathematicians shy away from. | |
| Understanding Lean: Code as Mathematical Proof | 3 | 6 | 1 | 1 | The host prompts for a definition of Lean and asks how Axiom compares to OpenAI and DeepMind's IMO systems. The guest educates the host on Curry-Howard correspondence, the history of automated theorem proving, and AlphaGeometry. | |
| Informal Natural Language vs. Formal Machine Code in Math | 4 | 5 | 2 | 2 | The host attempts to categorize DeepMind as informal brute-force compared to Axiom's neuro-symbolic approach. The guest corrects the host's premise, pointing out that AlphaProof was in fact a formal theorem proving system. | |
| Verification Beyond Final Answers: RLVR in Proof-Based Math | 3 | 5 | 1 | 1 | The host asks about reinforcement learning with verifiable rewards (RLVR). The guest explains the flaw of numerical answer verification (reward hacking) versus step-by-step proof verification in adult research math. | |
| Eliminating Hallucinations: Expanding Verification to Code and Beyond | 3 | 4 | 1 | 1 | The host asks whether each domain requires its own Lean equivalent as Axiom expands beyond math. The guest details why strongly typed languages like Rust are structurally much closer to Lean than natural language. | |
| The Future Roadmap: From Math Verification to AI for Science | 3 | 4 | 1 | 1 | The host asks where the threshold lies between scientific code and general text before pivoting to the guest's early life in China. The guest draws an analogy between her experience at Ross Math Camp and Axiom Prover's skill library. | |
| Academic Journey: MIT, Rhodes Scholarship, and Neuroscience at UCL Gatsby | 2 | 3 | 0 | 0 | The host asks about the guest's unique background combining MIT, a Rhodes scholarship in neuroscience at UCL Gatsby, and a JD/PhD at Stanford. The guest explains her interdisciplinary path collaboratively. | |
| Axiom Architecture: Conjecture, Prover, and Knowledge Base | 3 | 5 | 1 | 1 | The host asks about product architecture, LLM usage, synthetic data, and compute requirements. The guest uses a ocean sailing metaphor to explain conjecture, prover, and knowledge base before detailing formal synthetic data fuzzing. | |
| Product Maturity, Research Frontiers, and Expert Hiring | 3 | 3 | 0 | 0 | The host inquires about current product maturity and the split between research and engineering. The guest explains how the team pushes research frontiers while selectively hardening capabilities. | |
| Scaling Reasoning Complexity: From Putnam Problems to Frontier Math | 3 | 5 | 1 | 1 | The host asks if there is an Everest in math for AI to tackle. The guest educates the host on journal publication tiers and details the jump in reasoning tree complexity from 40 nodes to thousands of nodes. | |
| Scaling Up vs. Scaling Out Across Mathematical and Physical Domains | 4 | 5 | 2 | 2 | The host asks if AI can win a Fields Medal. The guest reframes the question, stating that mathematicians aim for the shortlist rather than the medal itself due to human factors, prompting the host to press on why. | |
| Catalyzing a Mathematical Renaissance and Scientific Breakthroughs | 2 | 4 | 1 | 0 | The host asks about AI generating groundbreaking scientific discoveries. The guest outlines shifting from a scarcity mindset around elite mathematical human labor to an era of AI-driven mathematical abundance. | |
| Theoretical Physics, 'Last Mile' Precision, and Discovery Cycles | 2 | 4 | 1 | 1 | The host prompts on timelines to reach this world. The guest discusses the value of last-mile precision and contrasts her view that math is code with CTO Shubo's view that code is math. | |
| Code as Math: High-Stakes Systems and Formal Verification | 1 | 4 | 0 | 0 | The host listens as the guest explains high-stakes industrial applications like GPU chip design and aerospace where partial verification is useless and formal guarantees are required. | |
| Mathematical Skepticism vs. Acceptance of Formal Verifiers | 3 | 4 | 1 | 1 | The host asks whether the broader math community fears replacement by AI. The guest explains that mathematicians reject uncheckable natural language LLM proofs but welcome Lean formal certificates. | |
| Evaluating Talent and Hiring for Taste | 3 | 3 | 0 | 0 | The host asks intelligent questions about evaluating talent, hiring for taste, and assembling world-class founding mathematicians. The guest shares her whale-frequency metaphor for recruiting top talent. | |
| Concluding the Interview and Expressing Gratitude | 0 | 0 | 0 | 0 | Standard conversational sign-off and podcast outro; no technical exchange or pushback. |