Sep 1, 2026 · 1h 3m · a16z

Can AI Learn Mathematical Intuition?

Daniel Litt · 36m spoken Lisha Li · 21m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of The a16z Show, mathematician Daniel Litt discusses how frontier artificial intelligence is reshaping mathematical research, contrasting AI brute-force computation with human conceptual understanding and exploring the broader implications for academic incentives and education.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The host as informed peer 5.3 Guest teaching 5.1 Guest disagreement 1.8 The host pushing back 0.4
05100:0015:0030:0045:001:00:000:57–4:28 · The host as informed peer 5/10 The a16z Show Title Sequence Lisha Li introduces Daniel Litt and asks for his perspective as a practicing mathematician on recent AI milestones. Litt breaks down the Erdős unit distance problem counterexample and explains why bringing in classical techniques across fields was meaningful.4:28–7:43 · The host as informed peer 6/10 Deconstructing Mathematical Reasoning and Model Chains of Thought Li characterizes model reasoning as symbol manipulation that might feel inhuman, but Litt immediately counters that the output chain of thought looks remarkably human and recognizable to a practicing mathematician.7:43–12:11 · The host as informed peer 6/10 Natural Language Scaling, Claude vs. ChatGPT, and Intuition Limits Li highlights that frontier models scale natural language reasoning rather than Lean formal verification. Litt agrees and elaborates on how models excel at applying known techniques across papers but struggle with non-rigorous philosophical intuition.12:11–20:14 · The host as informed peer 5/10 Frontier Model Evolution and Theory of Mind in Math Explanations Li suggests GPT-5.6 exhibits a better theory of mind regarding what the user knows, but Litt bluntly remarks that he finds all current models equally bad at theory of mind. He then provides an extensive breakdown of open problem solving versus theory building in algebraic geometry.20:14–23:41 · The host as informed peer 5/10 Mathematical Philosophy: Aesthetic Pursuits vs. Conceptual Physics Li asks about mathematical aesthetics and beauty driving pure math research, but Litt pushes back against aesthetic motivations, stating he prefers to view math as conceptual physics aiming for fundamental understanding rather than artistic elegance.23:41–28:18 · The host as informed peer 5/10 Why True Conjectures and Theory Building Challenge AI Models Litt explains why true conjectures embedded in broad theoretical frameworks are much harder for AI than finding counterexamples, as they require entirely new techniques rather than applying existing machinery.28:18–34:05 · The host as informed peer 5/10 Conceptual Understanding vs. Brute-Force Grinding in Mathematics Litt shares a personal anecdote where his inability to stomach an ugly brute-force calculation forced him to find a deeper, more conceptual proof, whereas an LLM happily generated a ten-page grind without insight.34:05–40:47 · The host as informed peer 6/10 Academic Publishing Incentives, Paper Slop, and Mode Collapse Litt discusses the risk of academic prestige gaming through low-effort AI paper generation and mode-collapsed solutions. Li adds that models trained on identical literature risk losing the diverse intuition human researchers bring.40:47–47:13 · The host as informed peer 6/10 Preserving Human Agency, Education, and Mathematical Capital Litt argues that even under superhuman AI, society must preserve human mathematical agency and diverse exploration. Li brings in pedagogy and Hungarian math curricula, noting the importance of teaching foundational thinking early.47:13–49:21 · The host as informed peer 4/10 Public Interest in Math and Evaluating Non-Expert Slop Litt nuances his complaint about slop by distinguishing academic paper gaming from enthusiastic amateur attempts, framing broad public engagement with mathematics as a net positive.49:21–51:55 · The host as informed peer 5/10 Rank 30 Elliptic Curves and Evaluating Constructive Record Results Li brings up the recent Claude-assisted rank 30 elliptic curve construction. Litt contextualizes it as an impressive record computation rather than a conceptual breakthrough that alters field theory.51:55–55:09 · The host as informed peer 5/10 Short Proofs, Verification Constraints, and Structural Paper Checking Litt explains that models currently produce short proofs not out of aesthetic discipline, but because verifying long arguments exceeds model and human capabilities. He cites an unverified 800-page AI preprint as an example of current limits.55:09–59:25 · The host as informed peer 6/10 AI Evaluation Harnesses and Proof Verification Li references Cursor's SQLite benchmark harness to discuss agentic architecture. Litt explains how human mathematicians stress-test global argument structures for subtle contradictions, a capability current models lack.59:25–1:02:44 · The host as informed peer 5/10 Early Math Education and Parenting in the AI Era Li and Litt bond over early childhood math education and how instilling mathematical clarity remains essential regardless of future AI capabilities.0:57–4:28 · Guest teaching 5/10 The a16z Show Title Sequence Lisha Li introduces Daniel Litt and asks for his perspective as a practicing mathematician on recent AI milestones. Litt breaks down the Erdős unit distance problem counterexample and explains why bringing in classical techniques across fields was meaningful.4:28–7:43 · Guest teaching 6/10 Deconstructing Mathematical Reasoning and Model Chains of Thought Li characterizes model reasoning as symbol manipulation that might feel inhuman, but Litt immediately counters that the output chain of thought looks remarkably human and recognizable to a practicing mathematician.7:43–12:11 · Guest teaching 5/10 Natural Language Scaling, Claude vs. ChatGPT, and Intuition Limits Li highlights that frontier models scale natural language reasoning rather than Lean formal verification. Litt agrees and elaborates on how models excel at applying known techniques across papers but struggle with non-rigorous philosophical intuition.12:11–20:14 · Guest teaching 5/10 Frontier Model Evolution and Theory of Mind in Math Explanations Li suggests GPT-5.6 exhibits a better theory of mind regarding what the user knows, but Litt bluntly remarks that he finds all current models equally bad at theory of mind. He then provides an extensive breakdown of open problem solving versus theory building in algebraic geometry.20:14–23:41 · Guest teaching 6/10 Mathematical Philosophy: Aesthetic Pursuits vs. Conceptual Physics Li asks about mathematical aesthetics and beauty driving pure math research, but Litt pushes back against aesthetic motivations, stating he prefers to view math as conceptual physics aiming for fundamental understanding rather than artistic elegance.23:41–28:18 · Guest teaching 5/10 Why True Conjectures and Theory Building Challenge AI Models Litt explains why true conjectures embedded in broad theoretical frameworks are much harder for AI than finding counterexamples, as they require entirely new techniques rather than applying existing machinery.28:18–34:05 · Guest teaching 6/10 Conceptual Understanding vs. Brute-Force Grinding in Mathematics Litt shares a personal anecdote where his inability to stomach an ugly brute-force calculation forced him to find a deeper, more conceptual proof, whereas an LLM happily generated a ten-page grind without insight.34:05–40:47 · Guest teaching 5/10 Academic Publishing Incentives, Paper Slop, and Mode Collapse Litt discusses the risk of academic prestige gaming through low-effort AI paper generation and mode-collapsed solutions. Li adds that models trained on identical literature risk losing the diverse intuition human researchers bring.40:47–47:13 · Guest teaching 5/10 Preserving Human Agency, Education, and Mathematical Capital Litt argues that even under superhuman AI, society must preserve human mathematical agency and diverse exploration. Li brings in pedagogy and Hungarian math curricula, noting the importance of teaching foundational thinking early.47:13–49:21 · Guest teaching 4/10 Public Interest in Math and Evaluating Non-Expert Slop Litt nuances his complaint about slop by distinguishing academic paper gaming from enthusiastic amateur attempts, framing broad public engagement with mathematics as a net positive.49:21–51:55 · Guest teaching 5/10 Rank 30 Elliptic Curves and Evaluating Constructive Record Results Li brings up the recent Claude-assisted rank 30 elliptic curve construction. Litt contextualizes it as an impressive record computation rather than a conceptual breakthrough that alters field theory.51:55–55:09 · Guest teaching 6/10 Short Proofs, Verification Constraints, and Structural Paper Checking Litt explains that models currently produce short proofs not out of aesthetic discipline, but because verifying long arguments exceeds model and human capabilities. He cites an unverified 800-page AI preprint as an example of current limits.55:09–59:25 · Guest teaching 5/10 AI Evaluation Harnesses and Proof Verification Li references Cursor's SQLite benchmark harness to discuss agentic architecture. Litt explains how human mathematicians stress-test global argument structures for subtle contradictions, a capability current models lack.59:25–1:02:44 · Guest teaching 4/10 Early Math Education and Parenting in the AI Era Li and Litt bond over early childhood math education and how instilling mathematical clarity remains essential regardless of future AI capabilities.0:57–4:28 · Guest disagreement 1/10 The a16z Show Title Sequence Lisha Li introduces Daniel Litt and asks for his perspective as a practicing mathematician on recent AI milestones. Litt breaks down the Erdős unit distance problem counterexample and explains why bringing in classical techniques across fields was meaningful.4:28–7:43 · Guest disagreement 4/10 Deconstructing Mathematical Reasoning and Model Chains of Thought Li characterizes model reasoning as symbol manipulation that might feel inhuman, but Litt immediately counters that the output chain of thought looks remarkably human and recognizable to a practicing mathematician.7:43–12:11 · Guest disagreement 1/10 Natural Language Scaling, Claude vs. ChatGPT, and Intuition Limits Li highlights that frontier models scale natural language reasoning rather than Lean formal verification. Litt agrees and elaborates on how models excel at applying known techniques across papers but struggle with non-rigorous philosophical intuition.12:11–20:14 · Guest disagreement 3/10 Frontier Model Evolution and Theory of Mind in Math Explanations Li suggests GPT-5.6 exhibits a better theory of mind regarding what the user knows, but Litt bluntly remarks that he finds all current models equally bad at theory of mind. He then provides an extensive breakdown of open problem solving versus theory building in algebraic geometry.20:14–23:41 · Guest disagreement 4/10 Mathematical Philosophy: Aesthetic Pursuits vs. Conceptual Physics Li asks about mathematical aesthetics and beauty driving pure math research, but Litt pushes back against aesthetic motivations, stating he prefers to view math as conceptual physics aiming for fundamental understanding rather than artistic elegance.23:41–28:18 · Guest disagreement 2/10 Why True Conjectures and Theory Building Challenge AI Models Litt explains why true conjectures embedded in broad theoretical frameworks are much harder for AI than finding counterexamples, as they require entirely new techniques rather than applying existing machinery.28:18–34:05 · Guest disagreement 2/10 Conceptual Understanding vs. Brute-Force Grinding in Mathematics Litt shares a personal anecdote where his inability to stomach an ugly brute-force calculation forced him to find a deeper, more conceptual proof, whereas an LLM happily generated a ten-page grind without insight.34:05–40:47 · Guest disagreement 1/10 Academic Publishing Incentives, Paper Slop, and Mode Collapse Litt discusses the risk of academic prestige gaming through low-effort AI paper generation and mode-collapsed solutions. Li adds that models trained on identical literature risk losing the diverse intuition human researchers bring.40:47–47:13 · Guest disagreement 1/10 Preserving Human Agency, Education, and Mathematical Capital Litt argues that even under superhuman AI, society must preserve human mathematical agency and diverse exploration. Li brings in pedagogy and Hungarian math curricula, noting the importance of teaching foundational thinking early.47:13–49:21 · Guest disagreement 2/10 Public Interest in Math and Evaluating Non-Expert Slop Litt nuances his complaint about slop by distinguishing academic paper gaming from enthusiastic amateur attempts, framing broad public engagement with mathematics as a net positive.49:21–51:55 · Guest disagreement 1/10 Rank 30 Elliptic Curves and Evaluating Constructive Record Results Li brings up the recent Claude-assisted rank 30 elliptic curve construction. Litt contextualizes it as an impressive record computation rather than a conceptual breakthrough that alters field theory.51:55–55:09 · Guest disagreement 2/10 Short Proofs, Verification Constraints, and Structural Paper Checking Litt explains that models currently produce short proofs not out of aesthetic discipline, but because verifying long arguments exceeds model and human capabilities. He cites an unverified 800-page AI preprint as an example of current limits.55:09–59:25 · Guest disagreement 1/10 AI Evaluation Harnesses and Proof Verification Li references Cursor's SQLite benchmark harness to discuss agentic architecture. Litt explains how human mathematicians stress-test global argument structures for subtle contradictions, a capability current models lack.59:25–1:02:44 · Guest disagreement 0/10 Early Math Education and Parenting in the AI Era Li and Litt bond over early childhood math education and how instilling mathematical clarity remains essential regardless of future AI capabilities.0:57–4:28 · The host pushing back 0/10 The a16z Show Title Sequence Lisha Li introduces Daniel Litt and asks for his perspective as a practicing mathematician on recent AI milestones. Litt breaks down the Erdős unit distance problem counterexample and explains why bringing in classical techniques across fields was meaningful.4:28–7:43 · The host pushing back 1/10 Deconstructing Mathematical Reasoning and Model Chains of Thought Li characterizes model reasoning as symbol manipulation that might feel inhuman, but Litt immediately counters that the output chain of thought looks remarkably human and recognizable to a practicing mathematician.7:43–12:11 · The host pushing back 0/10 Natural Language Scaling, Claude vs. ChatGPT, and Intuition Limits Li highlights that frontier models scale natural language reasoning rather than Lean formal verification. Litt agrees and elaborates on how models excel at applying known techniques across papers but struggle with non-rigorous philosophical intuition.12:11–20:14 · The host pushing back 1/10 Frontier Model Evolution and Theory of Mind in Math Explanations Li suggests GPT-5.6 exhibits a better theory of mind regarding what the user knows, but Litt bluntly remarks that he finds all current models equally bad at theory of mind. He then provides an extensive breakdown of open problem solving versus theory building in algebraic geometry.20:14–23:41 · The host pushing back 2/10 Mathematical Philosophy: Aesthetic Pursuits vs. Conceptual Physics Li asks about mathematical aesthetics and beauty driving pure math research, but Litt pushes back against aesthetic motivations, stating he prefers to view math as conceptual physics aiming for fundamental understanding rather than artistic elegance.23:41–28:18 · The host pushing back 0/10 Why True Conjectures and Theory Building Challenge AI Models Litt explains why true conjectures embedded in broad theoretical frameworks are much harder for AI than finding counterexamples, as they require entirely new techniques rather than applying existing machinery.28:18–34:05 · The host pushing back 1/10 Conceptual Understanding vs. Brute-Force Grinding in Mathematics Litt shares a personal anecdote where his inability to stomach an ugly brute-force calculation forced him to find a deeper, more conceptual proof, whereas an LLM happily generated a ten-page grind without insight.34:05–40:47 · The host pushing back 0/10 Academic Publishing Incentives, Paper Slop, and Mode Collapse Litt discusses the risk of academic prestige gaming through low-effort AI paper generation and mode-collapsed solutions. Li adds that models trained on identical literature risk losing the diverse intuition human researchers bring.40:47–47:13 · The host pushing back 0/10 Preserving Human Agency, Education, and Mathematical Capital Litt argues that even under superhuman AI, society must preserve human mathematical agency and diverse exploration. Li brings in pedagogy and Hungarian math curricula, noting the importance of teaching foundational thinking early.47:13–49:21 · The host pushing back 0/10 Public Interest in Math and Evaluating Non-Expert Slop Litt nuances his complaint about slop by distinguishing academic paper gaming from enthusiastic amateur attempts, framing broad public engagement with mathematics as a net positive.49:21–51:55 · The host pushing back 0/10 Rank 30 Elliptic Curves and Evaluating Constructive Record Results Li brings up the recent Claude-assisted rank 30 elliptic curve construction. Litt contextualizes it as an impressive record computation rather than a conceptual breakthrough that alters field theory.51:55–55:09 · The host pushing back 0/10 Short Proofs, Verification Constraints, and Structural Paper Checking Litt explains that models currently produce short proofs not out of aesthetic discipline, but because verifying long arguments exceeds model and human capabilities. He cites an unverified 800-page AI preprint as an example of current limits.55:09–59:25 · The host pushing back 0/10 AI Evaluation Harnesses and Proof Verification Li references Cursor's SQLite benchmark harness to discuss agentic architecture. Litt explains how human mathematicians stress-test global argument structures for subtle contradictions, a capability current models lack.59:25–1:02:44 · The host pushing back 0/10 Early Math Education and Parenting in the AI Era Li and Litt bond over early childhood math education and how instilling mathematical clarity remains essential regardless of future AI capabilities.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 0% · guest 100%0:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%33:00 · the host 0% · guest 100%33:00 · the host 0% · guest 100%36:00 · the host 0% · guest 100%36:00 · the host 0% · guest 100%39:00 · the host 0% · guest 100%39:00 · the host 0% · guest 100%42:00 · the host 0% · guest 100%42:00 · the host 0% · guest 100%45:00 · the host 0% · guest 100%45:00 · the host 0% · guest 100%48:00 · the host 0% · guest 100%48:00 · the host 0% · guest 100%51:00 · the host 0% · guest 100%51:00 · the host 0% · guest 100%54:00 · the host 0% · guest 100%54:00 · the host 0% · guest 100%57:00 · the host 0% · guest 100%57:00 · the host 0% · guest 100%1:00:00 · the host 0% · guest 100%1:00:00 · the host 0% · guest 100%
Sharpest disagreement ▶ 6:19 Litt contradicts host on AI proofs being inhuman

Litt directly rejects Li's characterization of AI proof generation as an inhuman feat of symbol manipulation, asserting that model outputs closely mirror human mathematical chains of thought.

Hardest push from the host ▶ 21:22 Litt pushes back on aesthetic motivation

When Li frames mathematical activity as driven by aesthetic beauty and sociological preference, Litt explicitly pushes back against using beauty as a guiding metric, advocating a physics-like conceptual approach instead.

Biggest teaching moment ▶ 52:58 Litt clarifies why models output short proofs

Litt educates the host on the reality behind short AI proofs, explaining that models do not choose brevity for elegance, but rather because neither models nor humans can verify correctness on long-horizon outputs.

The host holds their own ▶ 55:30 Li draws parallels to software harness evaluations

Li demonstrates deep technical familiarity with AI capabilities by citing Cursor's long-horizon harness experiments on SQLite in Rust to contextualize LLM planning limits in complex domains.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
The a16z Show Title Sequence 5510 Lisha Li introduces Daniel Litt and asks for his perspective as a practicing mathematician on recent AI milestones. Litt breaks down the Erdős unit distance problem counterexample and explains why bringing in classical techniques across fields was meaningful.
Deconstructing Mathematical Reasoning and Model Chains of Thought 6641 Li characterizes model reasoning as symbol manipulation that might feel inhuman, but Litt immediately counters that the output chain of thought looks remarkably human and recognizable to a practicing mathematician.
Natural Language Scaling, Claude vs. ChatGPT, and Intuition Limits 6510 Li highlights that frontier models scale natural language reasoning rather than Lean formal verification. Litt agrees and elaborates on how models excel at applying known techniques across papers but struggle with non-rigorous philosophical intuition.
Frontier Model Evolution and Theory of Mind in Math Explanations 5531 Li suggests GPT-5.6 exhibits a better theory of mind regarding what the user knows, but Litt bluntly remarks that he finds all current models equally bad at theory of mind. He then provides an extensive breakdown of open problem solving versus theory building in algebraic geometry.
Mathematical Philosophy: Aesthetic Pursuits vs. Conceptual Physics 5642 Li asks about mathematical aesthetics and beauty driving pure math research, but Litt pushes back against aesthetic motivations, stating he prefers to view math as conceptual physics aiming for fundamental understanding rather than artistic elegance.
Why True Conjectures and Theory Building Challenge AI Models 5520 Litt explains why true conjectures embedded in broad theoretical frameworks are much harder for AI than finding counterexamples, as they require entirely new techniques rather than applying existing machinery.
Conceptual Understanding vs. Brute-Force Grinding in Mathematics 5621 Litt shares a personal anecdote where his inability to stomach an ugly brute-force calculation forced him to find a deeper, more conceptual proof, whereas an LLM happily generated a ten-page grind without insight.
Academic Publishing Incentives, Paper Slop, and Mode Collapse 6510 Litt discusses the risk of academic prestige gaming through low-effort AI paper generation and mode-collapsed solutions. Li adds that models trained on identical literature risk losing the diverse intuition human researchers bring.
Preserving Human Agency, Education, and Mathematical Capital 6510 Litt argues that even under superhuman AI, society must preserve human mathematical agency and diverse exploration. Li brings in pedagogy and Hungarian math curricula, noting the importance of teaching foundational thinking early.
Public Interest in Math and Evaluating Non-Expert Slop 4420 Litt nuances his complaint about slop by distinguishing academic paper gaming from enthusiastic amateur attempts, framing broad public engagement with mathematics as a net positive.
Rank 30 Elliptic Curves and Evaluating Constructive Record Results 5510 Li brings up the recent Claude-assisted rank 30 elliptic curve construction. Litt contextualizes it as an impressive record computation rather than a conceptual breakthrough that alters field theory.
Short Proofs, Verification Constraints, and Structural Paper Checking 5620 Litt explains that models currently produce short proofs not out of aesthetic discipline, but because verifying long arguments exceeds model and human capabilities. He cites an unverified 800-page AI preprint as an example of current limits.
AI Evaluation Harnesses and Proof Verification 6510 Li references Cursor's SQLite benchmark harness to discuss agentic architecture. Litt explains how human mathematicians stress-test global argument structures for subtle contradictions, a capability current models lack.
Early Math Education and Parenting in the AI Era 5400 Li and Litt bond over early childhood math education and how instilling mathematical clarity remains essential regardless of future AI capabilities.

Statements from this episode (23)

Opinion
Litt: Many AI math breakthroughs are merely 'last mile' completions of human work
“So I think some of the results we've seen have had kind of the flavor of, like, you know, you kind of take some known techniques and apply them in maybe a clever way, or you I don't know, they've kind of been some kind of results I would characterize as, like,…”
Daniel Litt Sep 1, 2026 ▶ 3:01
Assertion Supported
Litt: Mathematicians used AI Erdős proof ideas to solve other open conjectures
“So a bunch of mathematicians took those ideas and used them to find counterexamples to a bunch of other interesting open questions. So, for example, like the sum product conjecture over the real numbers.”
Daniel Litt Sep 1, 2026 ▶ 4:01
Opinion
Litt: AI mathematical proofs resemble human reasoning, not alien 'Move 37' leaps
“And I would say that's actually, like, kind of typical of most of the results that I've studied. Like, they don't seem inhuman at all. They seem absolutely like something a human mathematician could produce. And they're, like, typically understandable if, like…”
Daniel Litt Sep 1, 2026 ▶ 6:52
Opinion
Litt: AI excels at computations but lacks mathematical intuition
“There'll be things that, this isn't surprising, like, there'll be things that rely on the model's strengths, like their ability to, like, grind out a long computation or, like, you know, pull together kind of technical ideas from many areas or, like, maybe man…”
Daniel Litt Sep 1, 2026 ▶ 10:06
Assertion Not checkable as stated
Litt: Frontier AI models cannot autonomously perform mathematical theory building
“So like, I've tried to get both, both Fable and ChatGPT, 5.6 Sol, I guess, to do some kind of theory building, and it's like, they're not, they definitely are not good at it autonomously, at least with, like, whatever scaffolding I've set up. But with some hin…”
Daniel Litt Sep 1, 2026 ▶ 11:31
Prediction Not checkable as stated
Litt predicts AI might autonomously build mathematical theories within six months
“My experience is that, like, if they can do it with, like, a hundred bits of hints or whatever in six months, maybe they can do it without hints.”
Daniel Litt Sep 1, 2026 ▶ 11:52
Assertion Not checkable as stated
Litt: Claude was useless for research math until Opus 4.5 or 4.6
“One thing is that ChatGPD got better at math earlier. Yes. So, like, for a long time the Claude models were just, like, not useful for research math. And then I think maybe around Opus 4.5 or Opus 4.6, they, like, more or less caught up.”
Daniel Litt Sep 1, 2026 ▶ 12:43
Insight
Litt: Open mathematical problems serve as benchmarks measuring lack of understanding
“At least for me the point of an open problem is it's, like, supposed to measure your failure to understand something. So it's kind of like a benchmark.”
Daniel Litt Sep 1, 2026 ▶ 15:02
Disclosure
Litt: AI acts as a Google substitute without doing deep intellectual math work
“What I've found is that the projects that I have that kind of predate AI, like the projects I've been thinking about for three or four or five years it's just not that useful. Like, it's primarily kind of a substitute for Google or something. Like, I might use…”
Daniel Litt Sep 1, 2026 ▶ 19:03
Insight
Litt: Reinforcement learning struggles to reward intermediate mathematical theory building
“I think what is definitely true is that, like, the skill of, like, developing a theory or, like, building your understanding of some poorly understood object is, like, a fuzzier one. So it might be harder, you know I guess you can try, you can tell it, you kno…”
Daniel Litt Sep 1, 2026 ▶ 27:08
Insight
Litt: Human inability to brute-force calculations drives profound mathematical discoveries
“And in fact, I think it's, like, kind of, like, our inability to just grind is kind of important to our ability to make discoveries.”
Daniel Litt Sep 1, 2026 ▶ 29:34
Disclosure
Litt: AI models successfully proved lemmas in one of his published papers
“So I've, I have one paper out so far where the models were kind of useful. So they, like, proved a couple lemmas in the paper.”
Daniel Litt Sep 1, 2026 ▶ 29:41
Opinion
Litt: Academic hiring incentives push math postdocs to generate AI slop
“Right now, I think, like, the existing incentive structures for math research do not do not encourage people to do that. So, you know, right now, if you're, like, a postdoc on the market, you want to get a job maybe for the next couple years before the communi…”
Daniel Litt Sep 1, 2026 ▶ 36:19
Disclosure
Litt: Prompting AI generated three correct algebraic geometry papers in one hour
“Here's an experiment you can do, you can take codecs, you can say, go online, find five recent conjectures in algebraic geometry and prove them, and ok, I've run this experiment, and with some back and forth, I was able to, you know, in an hour, get like three…”
Daniel Litt Sep 1, 2026 ▶ 36:48
Assertion Supported
Litt: Multiple preprint papers have appeared with identical AI-generated proofs
“Like, sometimes, you know, we've seen examples where, like, three or four or five papers with the exact same proof of the exact same theorem have come out in, within a couple days of each other, which is clearly, you know, some situation where someone's playin…”
Daniel Litt Sep 1, 2026 ▶ 37:46
Opinion
Litt: Society will still need human mathematicians even if AI becomes superhuman
“So, like, let's suppose the models become, like, really robustly superhuman, like, even, like, we're not even adding, like, meaningful cognitive diversity. Like, I claim, like, still, actually, we still want human mathematicians.”
Daniel Litt Sep 1, 2026 ▶ 40:50
Insight
Litt: Cheaper, lower-quality AI outputs risk displacing high-quality human work
“Like, you have a new technology that's doing something a little bit worse than was previously done, but much cheaper, and so you get a lot of, suddenly, a lot of, like, low-quality outputs that are displacing previous high-quality outputs.”
Daniel Litt Sep 1, 2026 ▶ 46:50
Opinion
Litt: Amateur AI math submissions are a net positive showing public enthusiasm
“So there's definitely also, like, slop coming from non-experts, but that I kind of actually don't see as a net negative. Like, ok, there's a lot of, now there's a lot of, like, documents on the internet one might have to comb through to figure out if a problem…”
Daniel Litt Sep 1, 2026 ▶ 48:52
Insight
Litt: AI mathematical results can only be properly evaluated in retrospect
“One thing I always say about a model result is, like, you cannot evaluate it except in retrospect, and like, this is also true of human mathematics. Like, sometimes a problem we thought was really important or would require really deep new ideas does not, and …”
Daniel Litt Sep 1, 2026 ▶ 51:05
Insight
Litt: AI cannot produce long proofs due to limits in verifying correctness
“The reason they're not producing long, complicated proofs is that they cannot. Like the, just like the ability to check correctness is not yet there.”
Daniel Litt Sep 1, 2026 ▶ 52:57
Opinion
Litt: An 800-page AI-generated math proof on arXiv is definitely incorrect
“Someone recently posted Acclaimed proof of resolution of singularities and positive characteristic, which was 800 AI generated pages. It's like, definitely, I mean, I'm sorry, I haven't read it. I haven't done an error, but there's no way it's correct. Like, t…”
Daniel Litt Sep 1, 2026 ▶ 54:17
Insight
Litt: AI harnesses designed to elicit proofs often decrease reliability
“When you make a harness whose goal is to elicit a proof, I think it often decreases reliability, because you're just trying to produce output.”
Daniel Litt Sep 1, 2026 ▶ 56:03
Insight
Litt: Math education remains valuable for clear thinking despite advanced AI
“I think a lot of what we educate people for is, like, pretty robust changes in the nature of the world. Like I think the reason to learn math has always been, like, to think clearly and, like, better understand the world, and, like, presumably that's something…”
Daniel Litt Sep 1, 2026 ▶ 1:01:02
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.