Oct 13, 2025 · 50m · a16z
Will LLMs Get Us To AGI?
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of The a16z Podcast, Columbia University Professor Vishal Misra joins Martin Casado and Erik Torenberg to discuss formal mathematical models of large language models (LLMs). Misra argues that while current transformers excel at navigating existing Bayesian manifolds, achieving true Artificial General Intelligence (AGI) requires fundamental architectural breakthroughs capable of creating entirely new scientific paradigms.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The host holds 3.1% of the talking time here. How this is scored →
speaking balance: gold is the host, purple is the guest (3 minute bins)
Vishal forcefully rejects the validity of 'prompt engineering', describing it mockingly as mere 'prompt twiddling' compared to real engineering.
Hardest push from the host ▶ 13:07 Challenge on Aesthetics PremiseMartin directly challenges and playfully ridicules Vishal's premise that a web form interface was a significant enough issue to motivate major software creation.
Biggest teaching moment ▶ 8:47 Explanation of Chain of Thought MechanicsVishal uses a step-by-step arithmetic multiplication example to demonstrate to the hosts precisely why chain-of-thought prompting reduces entropy.
The host holds their own ▶ 2:12 Domain Expertise BreakdownMartin demonstrates deep theoretical understanding by articulating Vishal's complex mathematical manifold model before the guest even explains it.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The host as informed peer | Guest teaching | Guest disagreement | The host pushing back | Why |
|---|---|---|---|---|---|---|
| Defining AGI: Going Beyond Pre-Trained Science | 6 | 3 | 1 | 2 | Martin demonstrates an unusually high level of technical understanding by summarizing Vishal's geometric manifold theory in detail before asking for feedback. Vishal confirms the summary is accurate while lighthearted banter about H-index scores ensues. | |
| Next-Token Distributions, Bayesian Manifolds, and Entropy | 1 | 6 | 1 | 0 | Vishal monologues on next-token probability distributions, Bayesian manifolds, and Shannon entropy. The hosts mostly listen with passive acknowledgement words while the guest details information entropy concepts. | |
| Chain of Thought and Algorithmic Reasoning | 3 | 6 | 1 | 3 | Martin briefly interrupts to ask for the core takeaway of state space reduction on reasoning. Vishal educates the host on why chain-of-thought works by using a step-by-step arithmetic multiplication analogy. | |
| How Solving a Cricket Problem Led to RAG | 2 | 5 | 2 | 3 | Vishal recounts how frustration with Crickinfo web forms led him to invent RAG via natural language DSL translation in 2020. Martin jokingly challenges the premise of being so personally bothered by an old web form interface. | |
| Rapid LLM Evolution and the Current Capability Plateau | 2 | 4 | 1 | 1 | Erik asks framing questions about LLM development pace and potential plateaus. Vishal elaborates on how LLMs are plateauing into incremental upgrades similar to modern iPhone release cycles. | |
| The Matrix Abstraction and In-Context Learning | 5 | 6 | 1 | 2 | Martin actively collaborates by validating how in-context learning acts as new Bayesian evidence. Vishal details the mathematical sparse matrix model and explains why prompt completion uses uniform underlying mechanics. | |
| Theoretical Limits: Why LLMs Cannot Recursively Self-Improve | 4 | 7 | 2 | 3 | Martin brings up common multi-agent feedback loop arguments for recursive self-improvement. Vishal rejects this premise, educating the host on why models are restricted to inductive closure and cannot discover new physics or axioms without external data. | |
| Defining AGI: Creating New Manifolds vs. Scaling Data | 4 | 6 | 2 | 3 | Martin raises the embodied AI counterargument that giving models real-world sensors could generate new learning. Vishal rejects scaling data alone as insufficient, stating architectural leaps are required to generate new manifolds. | |
| Future AI Architectures Beyond Pure Language | 5 | 5 | 1 | 2 | The group speculates on alternative architectures beyond pure text, such as JEPA and internal mental simulations. Martin draws upon observational sign language studies to discuss whether language follows intelligence. | |
| Theory vs. Empiricism and the Critique of "Prompt Engineering" | 4 | 5 | 3 | 2 | Vishal criticizes extreme empiricism in machine learning, mockingly dismissing prompt engineering as non-rigorous twiddling. Martin agrees and compares ML's empirical focus to historic systems engineering paradigms. | |
| Autonomous Software Creation, Multimodal Manifolds, and TokenProbe | 5 | 4 | 1 | 2 | Martin synthesizes Vishal's manifold framework to re-frame Erik's question about AGI benchmarks. The conversation concludes on a friendly note highlighting Vishal's TokenProbe tool. |