Jun 3, 2026 · 1h 33m · latent-space

Scaling Past Informal AI - Carina Hong, Axiom Math

Carina Hong · 1h 7m spoken RJ Haneke · 14m spoken Brandon Anderson · 4m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Axiom Math CEO Carina Hong discusses how formal verification in Lean overcomes the scaling limits of informal AI models to enable compound superintelligence. She details Axiom's breakthrough architectures, commercial opportunities in hardware and software verification, and the advantages of dedicated startup execution over frontier tech giants.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 5.3 Guest teaching 5.7 Guest disagreement 2.5 The hosts pushing back 3.7
05100:0020:0040:001:00:001:20:001:29–5:05 · The hosts as informed peer 4/10 Axiom's $200M Series A and Horizontal Transfer Learning Brandon questions how Axiom's valuation is justified given standard math research budgets. Carina reframes math from a niche vertical into a horizontal reasoning substrate by drawing parallels to Anthropic's early coding focus.5:05–10:52 · The hosts as informed peer 5/10 Frontier Lab Dynamics Versus Dedicated Startup Execution RJ asks why frontier labs wouldn't dominate formal verification. Carina explains lab organizational shifts following AlphaProof and argues verified AI is about scaling brilliance rather than merely mitigating hallucinations.10:52–13:24 · The hosts as informed peer 5/10 Deconstructing Lean as a Formal Language and Functional Tool Brandon asks Carina to explain Lean for non-experts. Carina and RJ discuss the Curry-Howard correspondence, Turing completeness, and Lean's functional nature.13:24–15:32 · The hosts as informed peer 4/10 The Verified Generation TAM and Putnam Competition Validation Carina forcefully rejects the assumption that formal verification has a small TAM limited to safety-critical niches, asserting it covers all AI code generation. She cites Axiom's 120/120 Putnam score beating DeepSeek.15:32–19:23 · The hosts as informed peer 6/10 Axiom's Technical Architecture and Cross-Domain Mathematical Coverage Brandon challenges whether recursive search risks getting trapped in narrow mathematical domains due to distribution shift. Carina concedes topology and analysis lack underlying Lean definitions compared to algebra.19:24–23:28 · The hosts as informed peer 5/10 Tackling Combinatorics Through Open-Source Mathematical Discovery Brandon asks why IMO combinatorics proved difficult for AlphaProof. Carina explains the need for creative constructions and previews Axiom open-sourcing non-Lean mathematical discovery tools.23:28–28:28 · The hosts as informed peer 7/10 Navigating Computational Limits and Decomposing Complex Programs RJ raises Rice's theorem and computational undecidability regarding verifying all programs. Carina acknowledges theoretical limits while explaining Axiom's strategy of decomposing complex control flows into verifiable subproblems.28:29–32:28 · The hosts as informed peer 6/10 Code Verification on Verina and Reinforcement Learning with Strong Types Carina shares benchmark metrics on Verina, contrasting typed RL in Rust/Lean with informal Python RL. RJ presses on how one knows generated formal code matches human problem intent.32:28–37:06 · The hosts as informed peer 6/10 The Specification Problem and Limits of Auto-Formalization RJ questions how verification handles real-world software specifications like flight control systems. Carina concedes human specification remains an open challenge and highlights automated unit test generation and auto-formalization.37:06–42:23 · The hosts as informed peer 6/10 Scaling Axiom Prover Trees and the Future of Human Comprehension RJ raises LLM context window and computational bounds on gigantic Lean proof trees. Carina details tree scaling from 40 to 4,000 nodes and cyclic auto-informalization.42:24–45:46 · The hosts as informed peer 5/10 Mathematical Elegance, Proof Diversity, and Human Cognitive Training Brandon inquires about optimizing for mathematical elegance, and RJ debates whether omitting low-level proof training harms high-level cognitive taste. Carina relates pre-training in Olympiad math to transferable intuition.45:47–50:10 · The hosts as informed peer 6/10 Hardware Verification Moats and the Economics of Verified Systems RJ presses on the commercial thesis justifying Axiom's valuation, citing 1:4 verification ratios in ASIC design. Carina outlines hardware verification where partial correctness has zero value.50:10–53:30 · The hosts as informed peer 5/10 Why Informal Reasoning Fails to Scale to Superintelligence RJ challenges Carina on why pure RL scaling on frontier LLMs cannot solve Math AGI informally. Carina explicitly goes on record rejecting informal math scaling due to the impossibility of scaling human grading and LLM judges.53:30–58:53 · The hosts as informed peer 4/10 Carina Hong's Academic Journey and Axiom's Talent Flywheel RJ explores Carina's background across Oxford Gatsby computational neuroscience, Stanford Law, and math. Carina explains how legal argumentation and appellate litigation transfer to mathematical reasoning.58:54–1:03:24 · The hosts as informed peer 6/10 The Erdős Conjecture Incident and the Challenge of Mathematical Search RJ brings up the controversy where Axiom formalized previously solved Erdős problems. Carina candidly owns the mistake and explains the technical difficulty of mathematical literature search and retrieval.1:03:25–1:06:57 · The hosts as informed peer 5/10 Self-Improvement Engines and Startup Agility Over Big Tech Brandon asks about AlphaZero-style self-improvement from scratch. Carina explains why dedicated startups can sustain focus on formal math while frontier labs suffer organizational turnover.1:06:57–1:11:07 · The hosts as informed peer 6/10 Launching the Axiom Lean Engine (AXLE) and Collaborative Proving RJ recounts using Axiom's AXLE API inside Claude Code. Carina explains the metaprogramming toolset behind AXLE and how automated blueprints can streamline large collaborative proofs.1:11:08–1:15:19 · The hosts as informed peer 6/10 Verification-Driven RL Rewards and API Infrastructure for Frontier Labs RJ asks about reward mechanisms in RL. Carina pitches Axiom's API as verification infrastructure for frontier labs that want rigorous reasoning rewards without maintaining Lean pipelines.1:15:19–1:21:12 · The hosts as informed peer 4/10 The Founding Genesis of Axiom and the True Purpose of Verified AI Brandon asks why Carina left Stanford to found Axiom. Carina delivers her core manifesto that verified AI is about compounding superhuman brilliance rather than policing hallucinations in closed industries.1:21:12–1:26:05 · The hosts as informed peer 5/10 The Bridge Between Reasoning, AI for Science, and Recursive Improvement The hosts and Carina discuss broader AI for science and ecosystem bottlenecks. Carina warns against venture-driven market fragmentation and premature commercialization distracting from deep technical capabilities.1:29–5:05 · Guest teaching 5/10 Axiom's $200M Series A and Horizontal Transfer Learning Brandon questions how Axiom's valuation is justified given standard math research budgets. Carina reframes math from a niche vertical into a horizontal reasoning substrate by drawing parallels to Anthropic's early coding focus.5:05–10:52 · Guest teaching 6/10 Frontier Lab Dynamics Versus Dedicated Startup Execution RJ asks why frontier labs wouldn't dominate formal verification. Carina explains lab organizational shifts following AlphaProof and argues verified AI is about scaling brilliance rather than merely mitigating hallucinations.10:52–13:24 · Guest teaching 6/10 Deconstructing Lean as a Formal Language and Functional Tool Brandon asks Carina to explain Lean for non-experts. Carina and RJ discuss the Curry-Howard correspondence, Turing completeness, and Lean's functional nature.13:24–15:32 · Guest teaching 7/10 The Verified Generation TAM and Putnam Competition Validation Carina forcefully rejects the assumption that formal verification has a small TAM limited to safety-critical niches, asserting it covers all AI code generation. She cites Axiom's 120/120 Putnam score beating DeepSeek.15:32–19:23 · Guest teaching 6/10 Axiom's Technical Architecture and Cross-Domain Mathematical Coverage Brandon challenges whether recursive search risks getting trapped in narrow mathematical domains due to distribution shift. Carina concedes topology and analysis lack underlying Lean definitions compared to algebra.19:24–23:28 · Guest teaching 6/10 Tackling Combinatorics Through Open-Source Mathematical Discovery Brandon asks why IMO combinatorics proved difficult for AlphaProof. Carina explains the need for creative constructions and previews Axiom open-sourcing non-Lean mathematical discovery tools.23:28–28:28 · Guest teaching 6/10 Navigating Computational Limits and Decomposing Complex Programs RJ raises Rice's theorem and computational undecidability regarding verifying all programs. Carina acknowledges theoretical limits while explaining Axiom's strategy of decomposing complex control flows into verifiable subproblems.28:29–32:28 · Guest teaching 6/10 Code Verification on Verina and Reinforcement Learning with Strong Types Carina shares benchmark metrics on Verina, contrasting typed RL in Rust/Lean with informal Python RL. RJ presses on how one knows generated formal code matches human problem intent.32:28–37:06 · Guest teaching 7/10 The Specification Problem and Limits of Auto-Formalization RJ questions how verification handles real-world software specifications like flight control systems. Carina concedes human specification remains an open challenge and highlights automated unit test generation and auto-formalization.37:06–42:23 · Guest teaching 5/10 Scaling Axiom Prover Trees and the Future of Human Comprehension RJ raises LLM context window and computational bounds on gigantic Lean proof trees. Carina details tree scaling from 40 to 4,000 nodes and cyclic auto-informalization.42:24–45:46 · Guest teaching 5/10 Mathematical Elegance, Proof Diversity, and Human Cognitive Training Brandon inquires about optimizing for mathematical elegance, and RJ debates whether omitting low-level proof training harms high-level cognitive taste. Carina relates pre-training in Olympiad math to transferable intuition.45:47–50:10 · Guest teaching 5/10 Hardware Verification Moats and the Economics of Verified Systems RJ presses on the commercial thesis justifying Axiom's valuation, citing 1:4 verification ratios in ASIC design. Carina outlines hardware verification where partial correctness has zero value.50:10–53:30 · Guest teaching 6/10 Why Informal Reasoning Fails to Scale to Superintelligence RJ challenges Carina on why pure RL scaling on frontier LLMs cannot solve Math AGI informally. Carina explicitly goes on record rejecting informal math scaling due to the impossibility of scaling human grading and LLM judges.53:30–58:53 · Guest teaching 6/10 Carina Hong's Academic Journey and Axiom's Talent Flywheel RJ explores Carina's background across Oxford Gatsby computational neuroscience, Stanford Law, and math. Carina explains how legal argumentation and appellate litigation transfer to mathematical reasoning.58:54–1:03:24 · Guest teaching 5/10 The Erdős Conjecture Incident and the Challenge of Mathematical Search RJ brings up the controversy where Axiom formalized previously solved Erdős problems. Carina candidly owns the mistake and explains the technical difficulty of mathematical literature search and retrieval.1:03:25–1:06:57 · Guest teaching 5/10 Self-Improvement Engines and Startup Agility Over Big Tech Brandon asks about AlphaZero-style self-improvement from scratch. Carina explains why dedicated startups can sustain focus on formal math while frontier labs suffer organizational turnover.1:06:57–1:11:07 · Guest teaching 5/10 Launching the Axiom Lean Engine (AXLE) and Collaborative Proving RJ recounts using Axiom's AXLE API inside Claude Code. Carina explains the metaprogramming toolset behind AXLE and how automated blueprints can streamline large collaborative proofs.1:11:08–1:15:19 · Guest teaching 6/10 Verification-Driven RL Rewards and API Infrastructure for Frontier Labs RJ asks about reward mechanisms in RL. Carina pitches Axiom's API as verification infrastructure for frontier labs that want rigorous reasoning rewards without maintaining Lean pipelines.1:15:19–1:21:12 · Guest teaching 6/10 The Founding Genesis of Axiom and the True Purpose of Verified AI Brandon asks why Carina left Stanford to found Axiom. Carina delivers her core manifesto that verified AI is about compounding superhuman brilliance rather than policing hallucinations in closed industries.1:21:12–1:26:05 · Guest teaching 5/10 The Bridge Between Reasoning, AI for Science, and Recursive Improvement The hosts and Carina discuss broader AI for science and ecosystem bottlenecks. Carina warns against venture-driven market fragmentation and premature commercialization distracting from deep technical capabilities.1:29–5:05 · Guest disagreement 3/10 Axiom's $200M Series A and Horizontal Transfer Learning Brandon questions how Axiom's valuation is justified given standard math research budgets. Carina reframes math from a niche vertical into a horizontal reasoning substrate by drawing parallels to Anthropic's early coding focus.5:05–10:52 · Guest disagreement 2/10 Frontier Lab Dynamics Versus Dedicated Startup Execution RJ asks why frontier labs wouldn't dominate formal verification. Carina explains lab organizational shifts following AlphaProof and argues verified AI is about scaling brilliance rather than merely mitigating hallucinations.10:52–13:24 · Guest disagreement 1/10 Deconstructing Lean as a Formal Language and Functional Tool Brandon asks Carina to explain Lean for non-experts. Carina and RJ discuss the Curry-Howard correspondence, Turing completeness, and Lean's functional nature.13:24–15:32 · Guest disagreement 4/10 The Verified Generation TAM and Putnam Competition Validation Carina forcefully rejects the assumption that formal verification has a small TAM limited to safety-critical niches, asserting it covers all AI code generation. She cites Axiom's 120/120 Putnam score beating DeepSeek.15:32–19:23 · Guest disagreement 2/10 Axiom's Technical Architecture and Cross-Domain Mathematical Coverage Brandon challenges whether recursive search risks getting trapped in narrow mathematical domains due to distribution shift. Carina concedes topology and analysis lack underlying Lean definitions compared to algebra.19:24–23:28 · Guest disagreement 2/10 Tackling Combinatorics Through Open-Source Mathematical Discovery Brandon asks why IMO combinatorics proved difficult for AlphaProof. Carina explains the need for creative constructions and previews Axiom open-sourcing non-Lean mathematical discovery tools.23:28–28:28 · Guest disagreement 3/10 Navigating Computational Limits and Decomposing Complex Programs RJ raises Rice's theorem and computational undecidability regarding verifying all programs. Carina acknowledges theoretical limits while explaining Axiom's strategy of decomposing complex control flows into verifiable subproblems.28:29–32:28 · Guest disagreement 2/10 Code Verification on Verina and Reinforcement Learning with Strong Types Carina shares benchmark metrics on Verina, contrasting typed RL in Rust/Lean with informal Python RL. RJ presses on how one knows generated formal code matches human problem intent.32:28–37:06 · Guest disagreement 3/10 The Specification Problem and Limits of Auto-Formalization RJ questions how verification handles real-world software specifications like flight control systems. Carina concedes human specification remains an open challenge and highlights automated unit test generation and auto-formalization.37:06–42:23 · Guest disagreement 2/10 Scaling Axiom Prover Trees and the Future of Human Comprehension RJ raises LLM context window and computational bounds on gigantic Lean proof trees. Carina details tree scaling from 40 to 4,000 nodes and cyclic auto-informalization.42:24–45:46 · Guest disagreement 3/10 Mathematical Elegance, Proof Diversity, and Human Cognitive Training Brandon inquires about optimizing for mathematical elegance, and RJ debates whether omitting low-level proof training harms high-level cognitive taste. Carina relates pre-training in Olympiad math to transferable intuition.45:47–50:10 · Guest disagreement 2/10 Hardware Verification Moats and the Economics of Verified Systems RJ presses on the commercial thesis justifying Axiom's valuation, citing 1:4 verification ratios in ASIC design. Carina outlines hardware verification where partial correctness has zero value.50:10–53:30 · Guest disagreement 6/10 Why Informal Reasoning Fails to Scale to Superintelligence RJ challenges Carina on why pure RL scaling on frontier LLMs cannot solve Math AGI informally. Carina explicitly goes on record rejecting informal math scaling due to the impossibility of scaling human grading and LLM judges.53:30–58:53 · Guest disagreement 1/10 Carina Hong's Academic Journey and Axiom's Talent Flywheel RJ explores Carina's background across Oxford Gatsby computational neuroscience, Stanford Law, and math. Carina explains how legal argumentation and appellate litigation transfer to mathematical reasoning.58:54–1:03:24 · Guest disagreement 3/10 The Erdős Conjecture Incident and the Challenge of Mathematical Search RJ brings up the controversy where Axiom formalized previously solved Erdős problems. Carina candidly owns the mistake and explains the technical difficulty of mathematical literature search and retrieval.1:03:25–1:06:57 · Guest disagreement 2/10 Self-Improvement Engines and Startup Agility Over Big Tech Brandon asks about AlphaZero-style self-improvement from scratch. Carina explains why dedicated startups can sustain focus on formal math while frontier labs suffer organizational turnover.1:06:57–1:11:07 · Guest disagreement 1/10 Launching the Axiom Lean Engine (AXLE) and Collaborative Proving RJ recounts using Axiom's AXLE API inside Claude Code. Carina explains the metaprogramming toolset behind AXLE and how automated blueprints can streamline large collaborative proofs.1:11:08–1:15:19 · Guest disagreement 2/10 Verification-Driven RL Rewards and API Infrastructure for Frontier Labs RJ asks about reward mechanisms in RL. Carina pitches Axiom's API as verification infrastructure for frontier labs that want rigorous reasoning rewards without maintaining Lean pipelines.1:15:19–1:21:12 · Guest disagreement 3/10 The Founding Genesis of Axiom and the True Purpose of Verified AI Brandon asks why Carina left Stanford to found Axiom. Carina delivers her core manifesto that verified AI is about compounding superhuman brilliance rather than policing hallucinations in closed industries.1:21:12–1:26:05 · Guest disagreement 2/10 The Bridge Between Reasoning, AI for Science, and Recursive Improvement The hosts and Carina discuss broader AI for science and ecosystem bottlenecks. Carina warns against venture-driven market fragmentation and premature commercialization distracting from deep technical capabilities.1:29–5:05 · The hosts pushing back 4/10 Axiom's $200M Series A and Horizontal Transfer Learning Brandon questions how Axiom's valuation is justified given standard math research budgets. Carina reframes math from a niche vertical into a horizontal reasoning substrate by drawing parallels to Anthropic's early coding focus.5:05–10:52 · The hosts pushing back 3/10 Frontier Lab Dynamics Versus Dedicated Startup Execution RJ asks why frontier labs wouldn't dominate formal verification. Carina explains lab organizational shifts following AlphaProof and argues verified AI is about scaling brilliance rather than merely mitigating hallucinations.10:52–13:24 · The hosts pushing back 2/10 Deconstructing Lean as a Formal Language and Functional Tool Brandon asks Carina to explain Lean for non-experts. Carina and RJ discuss the Curry-Howard correspondence, Turing completeness, and Lean's functional nature.13:24–15:32 · The hosts pushing back 2/10 The Verified Generation TAM and Putnam Competition Validation Carina forcefully rejects the assumption that formal verification has a small TAM limited to safety-critical niches, asserting it covers all AI code generation. She cites Axiom's 120/120 Putnam score beating DeepSeek.15:32–19:23 · The hosts pushing back 5/10 Axiom's Technical Architecture and Cross-Domain Mathematical Coverage Brandon challenges whether recursive search risks getting trapped in narrow mathematical domains due to distribution shift. Carina concedes topology and analysis lack underlying Lean definitions compared to algebra.19:24–23:28 · The hosts pushing back 3/10 Tackling Combinatorics Through Open-Source Mathematical Discovery Brandon asks why IMO combinatorics proved difficult for AlphaProof. Carina explains the need for creative constructions and previews Axiom open-sourcing non-Lean mathematical discovery tools.23:28–28:28 · The hosts pushing back 6/10 Navigating Computational Limits and Decomposing Complex Programs RJ raises Rice's theorem and computational undecidability regarding verifying all programs. Carina acknowledges theoretical limits while explaining Axiom's strategy of decomposing complex control flows into verifiable subproblems.28:29–32:28 · The hosts pushing back 4/10 Code Verification on Verina and Reinforcement Learning with Strong Types Carina shares benchmark metrics on Verina, contrasting typed RL in Rust/Lean with informal Python RL. RJ presses on how one knows generated formal code matches human problem intent.32:28–37:06 · The hosts pushing back 5/10 The Specification Problem and Limits of Auto-Formalization RJ questions how verification handles real-world software specifications like flight control systems. Carina concedes human specification remains an open challenge and highlights automated unit test generation and auto-formalization.37:06–42:23 · The hosts pushing back 5/10 Scaling Axiom Prover Trees and the Future of Human Comprehension RJ raises LLM context window and computational bounds on gigantic Lean proof trees. Carina details tree scaling from 40 to 4,000 nodes and cyclic auto-informalization.42:24–45:46 · The hosts pushing back 4/10 Mathematical Elegance, Proof Diversity, and Human Cognitive Training Brandon inquires about optimizing for mathematical elegance, and RJ debates whether omitting low-level proof training harms high-level cognitive taste. Carina relates pre-training in Olympiad math to transferable intuition.45:47–50:10 · The hosts pushing back 4/10 Hardware Verification Moats and the Economics of Verified Systems RJ presses on the commercial thesis justifying Axiom's valuation, citing 1:4 verification ratios in ASIC design. Carina outlines hardware verification where partial correctness has zero value.50:10–53:30 · The hosts pushing back 6/10 Why Informal Reasoning Fails to Scale to Superintelligence RJ challenges Carina on why pure RL scaling on frontier LLMs cannot solve Math AGI informally. Carina explicitly goes on record rejecting informal math scaling due to the impossibility of scaling human grading and LLM judges.53:30–58:53 · The hosts pushing back 2/10 Carina Hong's Academic Journey and Axiom's Talent Flywheel RJ explores Carina's background across Oxford Gatsby computational neuroscience, Stanford Law, and math. Carina explains how legal argumentation and appellate litigation transfer to mathematical reasoning.58:54–1:03:24 · The hosts pushing back 5/10 The Erdős Conjecture Incident and the Challenge of Mathematical Search RJ brings up the controversy where Axiom formalized previously solved Erdős problems. Carina candidly owns the mistake and explains the technical difficulty of mathematical literature search and retrieval.1:03:25–1:06:57 · The hosts pushing back 3/10 Self-Improvement Engines and Startup Agility Over Big Tech Brandon asks about AlphaZero-style self-improvement from scratch. Carina explains why dedicated startups can sustain focus on formal math while frontier labs suffer organizational turnover.1:06:57–1:11:07 · The hosts pushing back 2/10 Launching the Axiom Lean Engine (AXLE) and Collaborative Proving RJ recounts using Axiom's AXLE API inside Claude Code. Carina explains the metaprogramming toolset behind AXLE and how automated blueprints can streamline large collaborative proofs.1:11:08–1:15:19 · The hosts pushing back 3/10 Verification-Driven RL Rewards and API Infrastructure for Frontier Labs RJ asks about reward mechanisms in RL. Carina pitches Axiom's API as verification infrastructure for frontier labs that want rigorous reasoning rewards without maintaining Lean pipelines.1:15:19–1:21:12 · The hosts pushing back 2/10 The Founding Genesis of Axiom and the True Purpose of Verified AI Brandon asks why Carina left Stanford to found Axiom. Carina delivers her core manifesto that verified AI is about compounding superhuman brilliance rather than policing hallucinations in closed industries.1:21:12–1:26:05 · The hosts pushing back 4/10 The Bridge Between Reasoning, AI for Science, and Recursive Improvement The hosts and Carina discuss broader AI for science and ecosystem bottlenecks. Carina warns against venture-driven market fragmentation and premature commercialization distracting from deep technical capabilities.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:12:00 · the hosts 0% · guest 100%1:12:00 · the hosts 0% · guest 100%1:15:00 · the hosts 0% · guest 100%1:15:00 · the hosts 0% · guest 100%1:18:00 · the hosts 0% · guest 100%1:18:00 · the hosts 0% · guest 100%1:21:00 · the hosts 0% · guest 100%1:21:00 · the hosts 0% · guest 100%1:24:00 · the hosts 0% · guest 100%1:24:00 · the hosts 0% · guest 100%1:27:00 · the hosts 0% · guest 100%1:27:00 · the hosts 0% · guest 100%1:30:00 · the hosts 0% · guest 100%1:30:00 · the hosts 0% · guest 100%1:33:00 · the hosts 0% · guest 0%1:33:00 · the hosts 0% · guest 0%
Sharpest disagreement ▶ 50:50 Rejecting informal LLM scaling to Math AGI

Carina explicitly goes on the record rejecting the consensus belief that informal RL scaling in frontier models can ever reach mathematical superintelligence.

Hardest push from the hosts ▶ 23:54 Challenging verification with Rice's theorem

RJ directly confronts the core company premise by citing fundamental computer science theory on the undecidability of program verification.

Biggest teaching moment ▶ 13:24 Reframing formal verification TAM and Putnam proof

Carina firmly disabuses the hosts of the idea that verification is a slow safety tax, demonstrating sample-efficiency gains and a perfect 120 Putnam score.

The host holds their own ▶ 47:45 Grounding verification economics in ASIC hardware ratios

RJ demonstrates domain knowledge by citing 1:4 engineering headcount ratios in chip verification to anchor Axiom's commercial TAM.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Axiom's $200M Series A and Horizontal Transfer Learning 4534 Brandon questions how Axiom's valuation is justified given standard math research budgets. Carina reframes math from a niche vertical into a horizontal reasoning substrate by drawing parallels to Anthropic's early coding focus.
Frontier Lab Dynamics Versus Dedicated Startup Execution 5623 RJ asks why frontier labs wouldn't dominate formal verification. Carina explains lab organizational shifts following AlphaProof and argues verified AI is about scaling brilliance rather than merely mitigating hallucinations.
Deconstructing Lean as a Formal Language and Functional Tool 5612 Brandon asks Carina to explain Lean for non-experts. Carina and RJ discuss the Curry-Howard correspondence, Turing completeness, and Lean's functional nature.
The Verified Generation TAM and Putnam Competition Validation 4742 Carina forcefully rejects the assumption that formal verification has a small TAM limited to safety-critical niches, asserting it covers all AI code generation. She cites Axiom's 120/120 Putnam score beating DeepSeek.
Axiom's Technical Architecture and Cross-Domain Mathematical Coverage 6625 Brandon challenges whether recursive search risks getting trapped in narrow mathematical domains due to distribution shift. Carina concedes topology and analysis lack underlying Lean definitions compared to algebra.
Tackling Combinatorics Through Open-Source Mathematical Discovery 5623 Brandon asks why IMO combinatorics proved difficult for AlphaProof. Carina explains the need for creative constructions and previews Axiom open-sourcing non-Lean mathematical discovery tools.
Navigating Computational Limits and Decomposing Complex Programs 7636 RJ raises Rice's theorem and computational undecidability regarding verifying all programs. Carina acknowledges theoretical limits while explaining Axiom's strategy of decomposing complex control flows into verifiable subproblems.
Code Verification on Verina and Reinforcement Learning with Strong Types 6624 Carina shares benchmark metrics on Verina, contrasting typed RL in Rust/Lean with informal Python RL. RJ presses on how one knows generated formal code matches human problem intent.
The Specification Problem and Limits of Auto-Formalization 6735 RJ questions how verification handles real-world software specifications like flight control systems. Carina concedes human specification remains an open challenge and highlights automated unit test generation and auto-formalization.
Scaling Axiom Prover Trees and the Future of Human Comprehension 6525 RJ raises LLM context window and computational bounds on gigantic Lean proof trees. Carina details tree scaling from 40 to 4,000 nodes and cyclic auto-informalization.
Mathematical Elegance, Proof Diversity, and Human Cognitive Training 5534 Brandon inquires about optimizing for mathematical elegance, and RJ debates whether omitting low-level proof training harms high-level cognitive taste. Carina relates pre-training in Olympiad math to transferable intuition.
Hardware Verification Moats and the Economics of Verified Systems 6524 RJ presses on the commercial thesis justifying Axiom's valuation, citing 1:4 verification ratios in ASIC design. Carina outlines hardware verification where partial correctness has zero value.
Why Informal Reasoning Fails to Scale to Superintelligence 5666 RJ challenges Carina on why pure RL scaling on frontier LLMs cannot solve Math AGI informally. Carina explicitly goes on record rejecting informal math scaling due to the impossibility of scaling human grading and LLM judges.
Carina Hong's Academic Journey and Axiom's Talent Flywheel 4612 RJ explores Carina's background across Oxford Gatsby computational neuroscience, Stanford Law, and math. Carina explains how legal argumentation and appellate litigation transfer to mathematical reasoning.
The Erdős Conjecture Incident and the Challenge of Mathematical Search 6535 RJ brings up the controversy where Axiom formalized previously solved Erdős problems. Carina candidly owns the mistake and explains the technical difficulty of mathematical literature search and retrieval.
Self-Improvement Engines and Startup Agility Over Big Tech 5523 Brandon asks about AlphaZero-style self-improvement from scratch. Carina explains why dedicated startups can sustain focus on formal math while frontier labs suffer organizational turnover.
Launching the Axiom Lean Engine (AXLE) and Collaborative Proving 6512 RJ recounts using Axiom's AXLE API inside Claude Code. Carina explains the metaprogramming toolset behind AXLE and how automated blueprints can streamline large collaborative proofs.
Verification-Driven RL Rewards and API Infrastructure for Frontier Labs 6623 RJ asks about reward mechanisms in RL. Carina pitches Axiom's API as verification infrastructure for frontier labs that want rigorous reasoning rewards without maintaining Lean pipelines.
The Founding Genesis of Axiom and the True Purpose of Verified AI 4632 Brandon asks why Carina left Stanford to found Axiom. Carina delivers her core manifesto that verified AI is about compounding superhuman brilliance rather than policing hallucinations in closed industries.
The Bridge Between Reasoning, AI for Science, and Recursive Improvement 5524 The hosts and Carina discuss broader AI for science and ecosystem bottlenecks. Carina warns against venture-driven market fragmentation and premature commercialization distracting from deep technical capabilities.

Statements from this episode (37)

Disclosure
Hong: Axiom Math is 7-8 months old with about 30 employees
“We are like a seven, eight months old company, so it definitely means a lot to us. It's a really cool milestone. We're currently about like 30 people now, right?”
Carina Hong Jun 3, 2026 ▶ 2:15
Insight
Hong: Structured and formal data enables broad horizontal transfer learning
“If you have more structured and formal data, it's going to be a lot more horizontal than the specific vertical we are tackling.”
Carina Hong Jun 3, 2026 ▶ 3:58
Assertion Supported
Hong: AI Solved All Non-Combinatorics IMO Problems Across 2024 and 2025
“Across 24 and 25, AI models could solve all the problems that are not combinatorics.”
Carina Hong Jun 3, 2026 ▶ 5:56
Assertion Not checkable as stated
Hong: DeepMind's Formal Math Slowdown Post-AlphaProof Was Non-Technical
“After AlphaProof, kind of like, we didn't see a lot of the formal math you know, results or kind of progress from Google DeepMind, and that's actually because of reasons that are not necessarily technical.”
Carina Hong Jun 3, 2026 ▶ 6:18
Assertion Not checkable as stated
Hong: Competitor's AI demo can be solved entirely by Lean's grind tactic
“We're talking about, for example, the grind tactic in Lean. It can currently handle a lot of mass proofs, like, at a very low level. And this is pretty shocking because I have seen, you know, actually another company working in the same space, like, you know, …”
Carina Hong Jun 3, 2026 ▶ 10:27
Insight
Hong: Formal verification in AI is about scaling superintelligence, not bug fixes
“It is not about, like, formal verification or verified AI to us. It's not just about handling or, like, kicking out the lousiness, the hallucinations, the mistakes. It's about scaling brilliance. It's about super intelligence.”
Carina Hong Jun 3, 2026 ▶ 13:02
Opinion
Hong: Formal verification TAM covers all AI-generated code, not niche applications
“No, that's not the TEM. The TEM is all code. The TEM is a right of first refusal on all AI-generated code. Like, right of first refusal, meaning, you know, you get to choose whether you want to verify it.”
Carina Hong Jun 3, 2026 ▶ 13:39
Insight
Hong: Scaling inference for formal math has almost no wall
“I think that we found scaling inference to have almost no wall recursively decomposing you know, approved goal into many sub goals and then learning to backtrack as well.”
Carina Hong Jun 3, 2026 ▶ 16:49
Assertion Supported
Hong: Axiom Math has solved open research problems across math subfields
“We have good performance, you know, having solved open research questions and number theory, commutative algebra, algebraic geometry, some discrete math that come into Rx and probability.”
Carina Hong Jun 3, 2026 ▶ 19:14
Prediction Not checkable as stated
Hong: Lean-Based Systems Will Struggle With Highly Creative Combinatorics
“I think a Lean-based system will struggle in those very creative places, which is why we at Axiom actually also invest on something called mathematical discovery.”
Carina Hong Jun 3, 2026 ▶ 20:25
Disclosure
Hong: Axiom Math to open source two mathematical discovery codebases
“We have some major news in the coming weeks, basically open sourcing entire code bases of mathematical discovery coming up.”
Carina Hong Jun 3, 2026 ▶ 20:38
Assertion Partly supported
Hong: Axiom's Francois Charton Solved Decades-Old Math Conjectures
“Francois Charton, a member of technical staff at Axiom, and he previously have done Patent Boost and End-to-end, you know, settle this proof, a thirty-year-old conjecture by finding a counterexample found the solution to a one-hundred-and-thirty-year-old probl…”
Carina Hong Jun 3, 2026 ▶ 21:57
Assertion Supported
Hong: Harmonic's Aristotle Verified an Erdős Problem Proof Found by GPT
“In fact, like, you know, GPT found a proof to an unsolved Erdos problem, and our competitor Harmonic, you know, Aristotle you know, verified it.”
Carina Hong Jun 3, 2026 ▶ 25:33
Assertion Open · timeframe Jun 2029
Hong: Axiom's unmodified Putnam system achieved 99% on Verina benchmark
“And we actually recently, with no modification to the Putnam system, we saw a 99% out of the 189 problems, we saw a 187, we missed only two code-wisp-proof.”
Carina Hong Jun 3, 2026 ▶ 29:46
Insight
Hong: Lean and Rust yield superior reinforcement learning convergence over Python
“If you want proof to be informal math, It's very annoying, because then that's, like, just makes objective function. Your code is something like Python, your proof is, say, natural language, math proof. You will not have very strong RL kind of performance, rig…”
Carina Hong Jun 3, 2026 ▶ 30:00
Prediction Not checkable as stated
Hong: Future coding will rely on automated test-generated specifications
“I think this is the future of coding. Yes, I think this is the future of coding. And I think this is where, you know, this is where I think even if we are supposed, like given the assumption that everything can be formally verified, you know, like studying sor…”
Carina Hong Jun 3, 2026 ▶ 34:09
Assertion Not checkable as stated
Hong: Each line of verified code currently takes 20 proof lines
“Currently, actually, you know, for each line of code written, there could be like 20 lines of proof.”
Carina Hong Jun 3, 2026 ▶ 36:18
Assertion Not checkable as stated
Hong: Axiom Prover has scaled proof trees from 40 to 4,000 nodes
“We have seen it scale from 40 notes to 4000 notes.”
Carina Hong Jun 3, 2026 ▶ 37:18
Disclosure
Hong: Axiom Prover is an ensemble of multiple post-trained models
“Action Prover is an ensemble system of multiple models that we do post-training.”
Carina Hong Jun 3, 2026 ▶ 37:25
Insight
Hong: Auto-informalizing Lean code into natural language is much easier than auto-formalizing
“Auto-informalization is a lot easier than auto-formalization minus the problem of no grounding, right?”
Carina Hong Jun 3, 2026 ▶ 39:28
Insight
Hong: Training AI for mathematical elegance is an alignment problem
“At one point we're gonna get to there because, you know, I think the conjecture will probably depend on what, you know, will probably depend on what we mean by taste, elegance. Feels like an alignment problem to me, you know? Like, you know, who gets to say wh…”
Carina Hong Jun 3, 2026 ▶ 43:05
Assertion Contradicted
Haneke: ASIC verification takes 3 to 4 times more resources than design
“My understanding is that the industry standard for design to verification in ASIC ASIC project is like one to three, one to four.”
RJ Haneke Jun 3, 2026 ▶ 47:52
Prediction Open · timeframe Jun 2029
Hong: Axiom Math will be worth $10 billion
“Because when we realize the dream, the company's gonna be worth ten billion.”
Carina Hong Jun 3, 2026 ▶ 50:19
Prediction Not checkable as stated
Hong: Informal math systems will not achieve math AGI
“I'm going to say on the record, we do not believe that an informal math system is going to be the math AGI solution.”
Carina Hong Jun 3, 2026 ▶ 50:41
Insight
Hong: Formalizing proofs into code yields superior AI performance
“We generally think that formal math and by sort of converting math proofs to programs to code give us much better performance.”
Carina Hong Jun 3, 2026 ▶ 51:37
Assertion Supported
Hong: Axiom and Harmonic mistakenly claimed solved Erdős problems were new
“So actually what happened was our competitor, Harmonic, decided to publicize that they have solved unsolved problems, Erdos number one two four and four 81, and then we trusted their literature review, believing that these problems are really, truly unsolved. …”
Carina Hong Jun 3, 2026 ▶ 59:19
Insight
Hong: Accumulating synthetic AI data is not a moat, just buffer
“I think everyone is trying to accumulate like a data, which is not a mode. It's just time and time mode. It's all about like, you know, whether you can execute fast enough to make sure that you have like a certain buffer because of say your data set, you know,…”
Carina Hong Jun 3, 2026 ▶ 1:03:05
Assertion Supported
Hong: All OpenAI formal math researchers have left the company
“No, no, they all left.”
Carina Hong Jun 3, 2026 ▶ 1:05:45
Disclosure
Hong: Axiom Math released AXLE, a Lean proof validation toolkit
“So we just released AXLE, A-X-L-E, stands for Axiom Lean Engine. And it's really a set of, kind of, proof validation and manipulation tools that are built for Lean in the language of Lean. So it's a bunch of metaprogramming tools.”
Carina Hong Jun 3, 2026 ▶ 1:07:18
Disclosure
Hong: Axiom's AXLE toolkit contains 14 Lean tools including VerifyProof
“And Excel is currently, I think, 14 like, such tools starting from Verify Proof, which is the sort of, To make sure that there's not, nothing weird you know, going on, like no, no sort of cheating by lean code.”
Carina Hong Jun 3, 2026 ▶ 1:07:59
Assertion Not checkable as stated
Hong: Claude plus AXLE is a go-to setup in Lean community
“And we have seen also, we have heard from a lot of the people that Claude plus Axel is kind of their go-to setup for now.”
Carina Hong Jun 3, 2026 ▶ 1:09:03
Prediction Not checkable as stated
Hong: Auto-generated blueprints will be the key bottleneck in formal math
“I think auto-generated Blueprint is going to be a technical bottleneck that many people are trying to solve around the same time.”
Carina Hong Jun 3, 2026 ▶ 1:10:56
What-if
Hong: Axiom could not have solved Putnam problems in time without AXLE
“And without it, we couldn't have solved it with I think the eight problems within the time limit. Definitely not, not within the time limit.”
Carina Hong Jun 3, 2026 ▶ 1:12:59
Assertion Contradicted
Hong: DeepSeek dissolved its formal reasoning team over strategic shift
“And we have since, for example, Deep Seek All right. Like originally having a formal team and then later dissolve that team because of strategic direction change.”
Carina Hong Jun 3, 2026 ▶ 1:14:16
Disclosure
Hong: Verification is the best first commercial market for Axiom
“The DNA of the company is math. We think that verification is the best first market.”
Carina Hong Jun 3, 2026 ▶ 1:23:54
Prediction Not checkable as stated
Hong: AI for Math Will Fragment as Axiom and Harmonic Lead
“I expect fragmentation to start to happen as Axiom and Harmonic establish category leadership.”
Carina Hong Jun 3, 2026 ▶ 1:30:39
Opinion
Hong: Commercial Pressure Risks Distracting AI Math Startups From Core Capability
“Potentially trying to prove commercial value is going to distract significantly from the core capability improvement.”
Carina Hong Jun 3, 2026 ▶ 1:32:28
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.