Oct 13, 2025 · 50m · a16z

Will LLMs Get Us To AGI?

Vishal Misra · 29m spoken Martin Casado · 11m spoken Erik Torenberg · 1m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of The a16z Podcast, Columbia University Professor Vishal Misra joins Martin Casado and Erik Torenberg to discuss formal mathematical models of large language models (LLMs). Misra argues that while current transformers excel at navigating existing Bayesian manifolds, achieving true Artificial General Intelligence (AGI) requires fundamental architectural breakthroughs capable of creating entirely new scientific paradigms.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The host holds 3.1% of the talking time here. How this is scored →

The host as informed peer 3.7 Guest teaching 5.2 Guest disagreement 1.4 The host pushing back 2.1
05100:0015:0030:0045:000:00–3:21 · The host as informed peer 6/10 Defining AGI: Going Beyond Pre-Trained Science Martin demonstrates an unusually high level of technical understanding by summarizing Vishal's geometric manifold theory in detail before asking for feedback. Vishal confirms the summary is accurate while lighthearted banter about H-index scores ensues.3:21–7:41 · The host as informed peer 1/10 Next-Token Distributions, Bayesian Manifolds, and Entropy Vishal monologues on next-token probability distributions, Bayesian manifolds, and Shannon entropy. The hosts mostly listen with passive acknowledgement words while the guest details information entropy concepts.7:41–10:40 · The host as informed peer 3/10 Chain of Thought and Algorithmic Reasoning Martin briefly interrupts to ask for the core takeaway of state space reduction on reasoning. Vishal educates the host on why chain-of-thought works by using a step-by-step arithmetic multiplication analogy.10:40–16:44 · The host as informed peer 2/10 How Solving a Cricket Problem Led to RAG Vishal recounts how frustration with Crickinfo web forms led him to invent RAG via natural language DSL translation in 2020. Martin jokingly challenges the premise of being so personally bothered by an old web form interface.16:44–19:22 · The host as informed peer 2/10 Rapid LLM Evolution and the Current Capability Plateau Erik asks framing questions about LLM development pace and potential plateaus. Vishal elaborates on how LLMs are plateauing into incremental upgrades similar to modern iPhone release cycles.19:22–28:50 · The host as informed peer 5/10 The Matrix Abstraction and In-Context Learning Martin actively collaborates by validating how in-context learning acts as new Bayesian evidence. Vishal details the mathematical sparse matrix model and explains why prompt completion uses uniform underlying mechanics.28:50–33:54 · The host as informed peer 4/10 Theoretical Limits: Why LLMs Cannot Recursively Self-Improve Martin brings up common multi-agent feedback loop arguments for recursive self-improvement. Vishal rejects this premise, educating the host on why models are restricted to inductive closure and cannot discover new physics or axioms without external data.33:54–36:57 · The host as informed peer 4/10 Defining AGI: Creating New Manifolds vs. Scaling Data Martin raises the embodied AI counterargument that giving models real-world sensors could generate new learning. Vishal rejects scaling data alone as insufficient, stating architectural leaps are required to generate new manifolds.36:57–42:10 · The host as informed peer 5/10 Future AI Architectures Beyond Pure Language The group speculates on alternative architectures beyond pure text, such as JEPA and internal mental simulations. Martin draws upon observational sign language studies to discuss whether language follows intelligence.42:10–44:58 · The host as informed peer 4/10 Theory vs. Empiricism and the Critique of "Prompt Engineering" Vishal criticizes extreme empiricism in machine learning, mockingly dismissing prompt engineering as non-rigorous twiddling. Martin agrees and compares ML's empirical focus to historic systems engineering paradigms.44:58–50:16 · The host as informed peer 5/10 Autonomous Software Creation, Multimodal Manifolds, and TokenProbe Martin synthesizes Vishal's manifold framework to re-frame Erik's question about AGI benchmarks. The conversation concludes on a friendly note highlighting Vishal's TokenProbe tool.0:00–3:21 · Guest teaching 3/10 Defining AGI: Going Beyond Pre-Trained Science Martin demonstrates an unusually high level of technical understanding by summarizing Vishal's geometric manifold theory in detail before asking for feedback. Vishal confirms the summary is accurate while lighthearted banter about H-index scores ensues.3:21–7:41 · Guest teaching 6/10 Next-Token Distributions, Bayesian Manifolds, and Entropy Vishal monologues on next-token probability distributions, Bayesian manifolds, and Shannon entropy. The hosts mostly listen with passive acknowledgement words while the guest details information entropy concepts.7:41–10:40 · Guest teaching 6/10 Chain of Thought and Algorithmic Reasoning Martin briefly interrupts to ask for the core takeaway of state space reduction on reasoning. Vishal educates the host on why chain-of-thought works by using a step-by-step arithmetic multiplication analogy.10:40–16:44 · Guest teaching 5/10 How Solving a Cricket Problem Led to RAG Vishal recounts how frustration with Crickinfo web forms led him to invent RAG via natural language DSL translation in 2020. Martin jokingly challenges the premise of being so personally bothered by an old web form interface.16:44–19:22 · Guest teaching 4/10 Rapid LLM Evolution and the Current Capability Plateau Erik asks framing questions about LLM development pace and potential plateaus. Vishal elaborates on how LLMs are plateauing into incremental upgrades similar to modern iPhone release cycles.19:22–28:50 · Guest teaching 6/10 The Matrix Abstraction and In-Context Learning Martin actively collaborates by validating how in-context learning acts as new Bayesian evidence. Vishal details the mathematical sparse matrix model and explains why prompt completion uses uniform underlying mechanics.28:50–33:54 · Guest teaching 7/10 Theoretical Limits: Why LLMs Cannot Recursively Self-Improve Martin brings up common multi-agent feedback loop arguments for recursive self-improvement. Vishal rejects this premise, educating the host on why models are restricted to inductive closure and cannot discover new physics or axioms without external data.33:54–36:57 · Guest teaching 6/10 Defining AGI: Creating New Manifolds vs. Scaling Data Martin raises the embodied AI counterargument that giving models real-world sensors could generate new learning. Vishal rejects scaling data alone as insufficient, stating architectural leaps are required to generate new manifolds.36:57–42:10 · Guest teaching 5/10 Future AI Architectures Beyond Pure Language The group speculates on alternative architectures beyond pure text, such as JEPA and internal mental simulations. Martin draws upon observational sign language studies to discuss whether language follows intelligence.42:10–44:58 · Guest teaching 5/10 Theory vs. Empiricism and the Critique of "Prompt Engineering" Vishal criticizes extreme empiricism in machine learning, mockingly dismissing prompt engineering as non-rigorous twiddling. Martin agrees and compares ML's empirical focus to historic systems engineering paradigms.44:58–50:16 · Guest teaching 4/10 Autonomous Software Creation, Multimodal Manifolds, and TokenProbe Martin synthesizes Vishal's manifold framework to re-frame Erik's question about AGI benchmarks. The conversation concludes on a friendly note highlighting Vishal's TokenProbe tool.0:00–3:21 · Guest disagreement 1/10 Defining AGI: Going Beyond Pre-Trained Science Martin demonstrates an unusually high level of technical understanding by summarizing Vishal's geometric manifold theory in detail before asking for feedback. Vishal confirms the summary is accurate while lighthearted banter about H-index scores ensues.3:21–7:41 · Guest disagreement 1/10 Next-Token Distributions, Bayesian Manifolds, and Entropy Vishal monologues on next-token probability distributions, Bayesian manifolds, and Shannon entropy. The hosts mostly listen with passive acknowledgement words while the guest details information entropy concepts.7:41–10:40 · Guest disagreement 1/10 Chain of Thought and Algorithmic Reasoning Martin briefly interrupts to ask for the core takeaway of state space reduction on reasoning. Vishal educates the host on why chain-of-thought works by using a step-by-step arithmetic multiplication analogy.10:40–16:44 · Guest disagreement 2/10 How Solving a Cricket Problem Led to RAG Vishal recounts how frustration with Crickinfo web forms led him to invent RAG via natural language DSL translation in 2020. Martin jokingly challenges the premise of being so personally bothered by an old web form interface.16:44–19:22 · Guest disagreement 1/10 Rapid LLM Evolution and the Current Capability Plateau Erik asks framing questions about LLM development pace and potential plateaus. Vishal elaborates on how LLMs are plateauing into incremental upgrades similar to modern iPhone release cycles.19:22–28:50 · Guest disagreement 1/10 The Matrix Abstraction and In-Context Learning Martin actively collaborates by validating how in-context learning acts as new Bayesian evidence. Vishal details the mathematical sparse matrix model and explains why prompt completion uses uniform underlying mechanics.28:50–33:54 · Guest disagreement 2/10 Theoretical Limits: Why LLMs Cannot Recursively Self-Improve Martin brings up common multi-agent feedback loop arguments for recursive self-improvement. Vishal rejects this premise, educating the host on why models are restricted to inductive closure and cannot discover new physics or axioms without external data.33:54–36:57 · Guest disagreement 2/10 Defining AGI: Creating New Manifolds vs. Scaling Data Martin raises the embodied AI counterargument that giving models real-world sensors could generate new learning. Vishal rejects scaling data alone as insufficient, stating architectural leaps are required to generate new manifolds.36:57–42:10 · Guest disagreement 1/10 Future AI Architectures Beyond Pure Language The group speculates on alternative architectures beyond pure text, such as JEPA and internal mental simulations. Martin draws upon observational sign language studies to discuss whether language follows intelligence.42:10–44:58 · Guest disagreement 3/10 Theory vs. Empiricism and the Critique of "Prompt Engineering" Vishal criticizes extreme empiricism in machine learning, mockingly dismissing prompt engineering as non-rigorous twiddling. Martin agrees and compares ML's empirical focus to historic systems engineering paradigms.44:58–50:16 · Guest disagreement 1/10 Autonomous Software Creation, Multimodal Manifolds, and TokenProbe Martin synthesizes Vishal's manifold framework to re-frame Erik's question about AGI benchmarks. The conversation concludes on a friendly note highlighting Vishal's TokenProbe tool.0:00–3:21 · The host pushing back 2/10 Defining AGI: Going Beyond Pre-Trained Science Martin demonstrates an unusually high level of technical understanding by summarizing Vishal's geometric manifold theory in detail before asking for feedback. Vishal confirms the summary is accurate while lighthearted banter about H-index scores ensues.3:21–7:41 · The host pushing back 0/10 Next-Token Distributions, Bayesian Manifolds, and Entropy Vishal monologues on next-token probability distributions, Bayesian manifolds, and Shannon entropy. The hosts mostly listen with passive acknowledgement words while the guest details information entropy concepts.7:41–10:40 · The host pushing back 3/10 Chain of Thought and Algorithmic Reasoning Martin briefly interrupts to ask for the core takeaway of state space reduction on reasoning. Vishal educates the host on why chain-of-thought works by using a step-by-step arithmetic multiplication analogy.10:40–16:44 · The host pushing back 3/10 How Solving a Cricket Problem Led to RAG Vishal recounts how frustration with Crickinfo web forms led him to invent RAG via natural language DSL translation in 2020. Martin jokingly challenges the premise of being so personally bothered by an old web form interface.16:44–19:22 · The host pushing back 1/10 Rapid LLM Evolution and the Current Capability Plateau Erik asks framing questions about LLM development pace and potential plateaus. Vishal elaborates on how LLMs are plateauing into incremental upgrades similar to modern iPhone release cycles.19:22–28:50 · The host pushing back 2/10 The Matrix Abstraction and In-Context Learning Martin actively collaborates by validating how in-context learning acts as new Bayesian evidence. Vishal details the mathematical sparse matrix model and explains why prompt completion uses uniform underlying mechanics.28:50–33:54 · The host pushing back 3/10 Theoretical Limits: Why LLMs Cannot Recursively Self-Improve Martin brings up common multi-agent feedback loop arguments for recursive self-improvement. Vishal rejects this premise, educating the host on why models are restricted to inductive closure and cannot discover new physics or axioms without external data.33:54–36:57 · The host pushing back 3/10 Defining AGI: Creating New Manifolds vs. Scaling Data Martin raises the embodied AI counterargument that giving models real-world sensors could generate new learning. Vishal rejects scaling data alone as insufficient, stating architectural leaps are required to generate new manifolds.36:57–42:10 · The host pushing back 2/10 Future AI Architectures Beyond Pure Language The group speculates on alternative architectures beyond pure text, such as JEPA and internal mental simulations. Martin draws upon observational sign language studies to discuss whether language follows intelligence.42:10–44:58 · The host pushing back 2/10 Theory vs. Empiricism and the Critique of "Prompt Engineering" Vishal criticizes extreme empiricism in machine learning, mockingly dismissing prompt engineering as non-rigorous twiddling. Martin agrees and compares ML's empirical focus to historic systems engineering paradigms.44:58–50:16 · The host pushing back 2/10 Autonomous Software Creation, Multimodal Manifolds, and TokenProbe Martin synthesizes Vishal's manifold framework to re-frame Erik's question about AGI benchmarks. The conversation concludes on a friendly note highlighting Vishal's TokenProbe tool.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 10.2% · guest 89.8%0:00 · the host 10.2% · guest 89.8%3:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%9:00 · the host 4.9% · guest 95.1%9:00 · the host 4.9% · guest 95.1%12:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%15:00 · the host 3.7% · guest 96.3%15:00 · the host 3.7% · guest 96.3%18:00 · the host 3.1% · guest 96.9%18:00 · the host 3.1% · guest 96.9%21:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%33:00 · the host 10.4% · guest 89.6%33:00 · the host 10.4% · guest 89.6%36:00 · the host 11% · guest 89%36:00 · the host 11% · guest 89%39:00 · the host 0% · guest 100%39:00 · the host 0% · guest 100%42:00 · the host 7.3% · guest 92.7%42:00 · the host 7.3% · guest 92.7%45:00 · the host 0% · guest 100%45:00 · the host 0% · guest 100%48:00 · the host 3.5% · guest 96.5%48:00 · the host 3.5% · guest 96.5%
Sharpest disagreement ▶ 43:54 Dismissal of Prompt Engineering

Vishal forcefully rejects the validity of 'prompt engineering', describing it mockingly as mere 'prompt twiddling' compared to real engineering.

Hardest push from the host ▶ 13:07 Challenge on Aesthetics Premise

Martin directly challenges and playfully ridicules Vishal's premise that a web form interface was a significant enough issue to motivate major software creation.

Biggest teaching moment ▶ 8:47 Explanation of Chain of Thought Mechanics

Vishal uses a step-by-step arithmetic multiplication example to demonstrate to the hosts precisely why chain-of-thought prompting reduces entropy.

The host holds their own ▶ 2:12 Domain Expertise Breakdown

Martin demonstrates deep theoretical understanding by articulating Vishal's complex mathematical manifold model before the guest even explains it.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
Defining AGI: Going Beyond Pre-Trained Science 6312 Martin demonstrates an unusually high level of technical understanding by summarizing Vishal's geometric manifold theory in detail before asking for feedback. Vishal confirms the summary is accurate while lighthearted banter about H-index scores ensues.
Next-Token Distributions, Bayesian Manifolds, and Entropy 1610 Vishal monologues on next-token probability distributions, Bayesian manifolds, and Shannon entropy. The hosts mostly listen with passive acknowledgement words while the guest details information entropy concepts.
Chain of Thought and Algorithmic Reasoning 3613 Martin briefly interrupts to ask for the core takeaway of state space reduction on reasoning. Vishal educates the host on why chain-of-thought works by using a step-by-step arithmetic multiplication analogy.
How Solving a Cricket Problem Led to RAG 2523 Vishal recounts how frustration with Crickinfo web forms led him to invent RAG via natural language DSL translation in 2020. Martin jokingly challenges the premise of being so personally bothered by an old web form interface.
Rapid LLM Evolution and the Current Capability Plateau 2411 Erik asks framing questions about LLM development pace and potential plateaus. Vishal elaborates on how LLMs are plateauing into incremental upgrades similar to modern iPhone release cycles.
The Matrix Abstraction and In-Context Learning 5612 Martin actively collaborates by validating how in-context learning acts as new Bayesian evidence. Vishal details the mathematical sparse matrix model and explains why prompt completion uses uniform underlying mechanics.
Theoretical Limits: Why LLMs Cannot Recursively Self-Improve 4723 Martin brings up common multi-agent feedback loop arguments for recursive self-improvement. Vishal rejects this premise, educating the host on why models are restricted to inductive closure and cannot discover new physics or axioms without external data.
Defining AGI: Creating New Manifolds vs. Scaling Data 4623 Martin raises the embodied AI counterargument that giving models real-world sensors could generate new learning. Vishal rejects scaling data alone as insufficient, stating architectural leaps are required to generate new manifolds.
Future AI Architectures Beyond Pure Language 5512 The group speculates on alternative architectures beyond pure text, such as JEPA and internal mental simulations. Martin draws upon observational sign language studies to discuss whether language follows intelligence.
Theory vs. Empiricism and the Critique of "Prompt Engineering" 4532 Vishal criticizes extreme empiricism in machine learning, mockingly dismissing prompt engineering as non-rigorous twiddling. Martin agrees and compares ML's empirical focus to historic systems engineering paradigms.
Autonomous Software Creation, Multimodal Manifolds, and TokenProbe 5412 Martin synthesizes Vishal's manifold framework to re-frame Erik's question about AGI benchmarks. The conversation concludes on a friendly note highlighting Vishal's TokenProbe tool.

Statements from this episode (13)

What-if
Misra: LLMs trained on pre-1915 physics could not discover relativity
“Any LLM that was trained on pre-nineteen-fifteen physics would never have come up with a theory of relativity.”
Vishal Misra Oct 13, 2025 ▶ 0:00
Opinion
Misra: AGI requires generating new science beyond training data
“AGI will be when we are able to create new science, new results, new math. When an AGI comes up with a theory of relativity, it has to go beyond what it has been trained on. To come up with new paradigms, new science. That's my definition of AGI.”
Vishal Misra Oct 13, 2025 ▶ 0:16
Insight
Casado: LLMs simplify reasoning by mapping complex spaces to geometric manifolds
“They reduce a very, very complex multidimensional space into Basically a geometric manifold that's a reduced state space. So it's reduced degrees of freedom, but you can actually predict where in the manifold the reasoning can move to.”
Martin Casado Oct 13, 2025 ▶ 2:14
Insight
Casado: Human reasoning operates by mapping reality onto lower-dimensional manifolds
“We as humans do the same thing as we take this very complex, heavy tailed, stochastic universe, and we reduce it to kind of this geometric manifold, and then When we reason, we just move along that manifold.”
Martin Casado Oct 13, 2025 ▶ 2:53
Assertion Not checkable as stated
Misra: LLMs hallucinate when straying from learned Bayesian manifolds
“And as long as the LLM is going in sort of traversing through these manifolds, it is confident. And it can produce something which is, which makes sense. The moment it sort of wears away from the manifold, then it starts hallucinating and start spotting nonsen…”
Vishal Misra Oct 13, 2025 ▶ 4:28
Insight
Misra: Adding context to prompts reduces LLM prediction entropy
“The moment you add more context. You make the prompt information rich. The prediction entropy reduces.”
Vishal Misra Oct 13, 2025 ▶ 7:28
Insight
Misra: Chain-of-thought prompting works by reducing LLM prediction entropy
“That's why chain of heart works. What happens with chain of thought is you ask the LLM to do something chain of thought. It starts breaking the problem into small steps. These steps it has seen in the past. It has been trained on maybe with some different numb…”
Vishal Misra Oct 13, 2025 ▶ 10:12
Assertion Supported
Misra: GPT-3 prompt matrix rows exceed atoms in all known galaxies
“If you just take just the old first generation GPT-III model, which had a context window of 2000 tokens and a vocabulary of 50,000 next tokens or 50,000 tokens, then the size of it, the number of rows in this matrix is more than the number of atoms across all …”
Vishal Misra Oct 13, 2025 ▶ 22:06
Insight
Misra: LLM output is bounded by the inductive closure of training data
“So you know another phrase that we have been using recently is, you know, the output of the LLM is the inductive closure of what it has been trained on.”
Vishal Misra Oct 13, 2025 ▶ 28:53
Assertion Not checkable as stated
Misra: Current LLM architectures cannot recursively self-improve into new paradigms
“That kind of self-improvement is not possible with these architectures. They can refine these. They can fill out these rows where the answer already exists.”
Vishal Misra Oct 13, 2025 ▶ 32:10
Insight
Misra: Adding data to LLMs cannot create new mathematical manifolds
“There has to be an architectural lead that is able to create these manifolds and just throwing new data will not do it. It'll just smoothen out the already existing manifolds.”
Vishal Misra Oct 13, 2025 ▶ 37:47
Insight
Misra: Pure language processing is insufficient for human-level intelligence
“Language is great, but language is not the answer. You know, when I'm looking at catching a ball that is coming to me, I'm mentally doing that simulation in my head. I'm not translating it to language to figure out where it'll land.”
Vishal Misra Oct 13, 2025 ▶ 39:22
Opinion
Misra: Prompt engineering is prompt twiddling, not true engineering
“One term I really dislike is prompt engineering. You know, engineering used to mean sending a man to the moon or providing five nines reliability. Prompt engineering is prompt twiddling.”
Vishal Misra Oct 13, 2025 ▶ 43:49
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.