Mar 17, 2026 · 46m · a16z

Why Scale Will Not Solve AGI | Vishal Misra - The a16z Show

Vishal Misra · 30m spoken Martin Casado · 8m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of The a16z Show, host Martin Casado interviews Columbia University Professor Vishal Misra about the mathematical mechanics of Large Language Models and why simple scaling will not achieve true Artificial General Intelligence (AGI). Misra argues that LLMs function as Bayesian inference engines over sparse probability matrices and contends that reaching AGI requires continuous synaptic plasticity and causal world models rather than larger compute clusters.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The host as informed peer 3.4 Guest teaching 5.1 Guest disagreement 1.9 The host pushing back 1.9
05100:0015:0030:0045:000:42–2:55 · The host as informed peer 2/10 Early AI Experiments and Origin of RAG Martin opens by welcoming back Vishal and framing his work on LLM mathematical modeling as top-tier. Vishal shares the origin story of using GPT-3 in 2020 to build an early RAG implementation for ESPN cricket stats.2:55–8:36 · The host as informed peer 3/10 The Matrix Abstraction of Language Models Martin asks targeted questions about row combinations and posterior distributions within Vishal's matrix model abstraction. Vishal explains how prompt rows generate vocabulary probability distributions and how huge context windows create astronomically large sparse matrices.8:36–15:02 · The host as informed peer 3/10 Explaining In-Context Learning as Bayesian Updating Martin notes that in-context learning was non-obvious to him and shares his experience using Vishal's domain-specific language (DSL). Vishal explains how showing few-shot examples continuously increases target token probabilities, demonstrating real-time Bayesian updating.15:02–23:02 · The host as informed peer 4/10 Proving Bayesian Behavior via the Bayesian Wind Tunnel Martin presses on why a second paper was necessary when the first paper already seemed empirically convincing. Vishal explains the community skepticism stemming from Bayesian versus frequentist debates and outlines how the Bayesian wind tunnel mathematically proved exact Bayesian posterior matching.23:02–30:19 · The host as informed peer 3/10 Human Cognition vs. LLM Mechanics and Consciousness Martin brings up public claims about potential LLM consciousness. Vishal forcefully rejects the premise, emphasizing that LLMs are frozen matrix multiplications driven by next-token optimization rather than plastic human brains driven by survival.30:19–38:54 · The host as informed peer 5/10 Why Scale Will Not Solve AGI and The Einstein Test Martin demonstrates strong domain insight by articulating how LLMs suffer from 'data gravity,' binding them to existing training consensus. Vishal uses Einstein's relativity theory to show why scaling existing correlation models cannot produce new Kolmogorov representations.38:54–46:25 · The host as informed peer 4/10 Causal Models, Knuth's Experiment, and Future Research Martin queries whether simulation mechanisms align with Kolmogorov complexity and asks about practical research directions. Vishal breaks down Donald Knuth's viral LLM experiment, showing that Knuth manually supplied the causal reasoning while the LLM performed association.0:42–2:55 · Guest teaching 3/10 Early AI Experiments and Origin of RAG Martin opens by welcoming back Vishal and framing his work on LLM mathematical modeling as top-tier. Vishal shares the origin story of using GPT-3 in 2020 to build an early RAG implementation for ESPN cricket stats.2:55–8:36 · Guest teaching 5/10 The Matrix Abstraction of Language Models Martin asks targeted questions about row combinations and posterior distributions within Vishal's matrix model abstraction. Vishal explains how prompt rows generate vocabulary probability distributions and how huge context windows create astronomically large sparse matrices.8:36–15:02 · Guest teaching 6/10 Explaining In-Context Learning as Bayesian Updating Martin notes that in-context learning was non-obvious to him and shares his experience using Vishal's domain-specific language (DSL). Vishal explains how showing few-shot examples continuously increases target token probabilities, demonstrating real-time Bayesian updating.15:02–23:02 · Guest teaching 6/10 Proving Bayesian Behavior via the Bayesian Wind Tunnel Martin presses on why a second paper was necessary when the first paper already seemed empirically convincing. Vishal explains the community skepticism stemming from Bayesian versus frequentist debates and outlines how the Bayesian wind tunnel mathematically proved exact Bayesian posterior matching.23:02–30:19 · Guest teaching 6/10 Human Cognition vs. LLM Mechanics and Consciousness Martin brings up public claims about potential LLM consciousness. Vishal forcefully rejects the premise, emphasizing that LLMs are frozen matrix multiplications driven by next-token optimization rather than plastic human brains driven by survival.30:19–38:54 · Guest teaching 5/10 Why Scale Will Not Solve AGI and The Einstein Test Martin demonstrates strong domain insight by articulating how LLMs suffer from 'data gravity,' binding them to existing training consensus. Vishal uses Einstein's relativity theory to show why scaling existing correlation models cannot produce new Kolmogorov representations.38:54–46:25 · Guest teaching 5/10 Causal Models, Knuth's Experiment, and Future Research Martin queries whether simulation mechanisms align with Kolmogorov complexity and asks about practical research directions. Vishal breaks down Donald Knuth's viral LLM experiment, showing that Knuth manually supplied the causal reasoning while the LLM performed association.0:42–2:55 · Guest disagreement 1/10 Early AI Experiments and Origin of RAG Martin opens by welcoming back Vishal and framing his work on LLM mathematical modeling as top-tier. Vishal shares the origin story of using GPT-3 in 2020 to build an early RAG implementation for ESPN cricket stats.2:55–8:36 · Guest disagreement 1/10 The Matrix Abstraction of Language Models Martin asks targeted questions about row combinations and posterior distributions within Vishal's matrix model abstraction. Vishal explains how prompt rows generate vocabulary probability distributions and how huge context windows create astronomically large sparse matrices.8:36–15:02 · Guest disagreement 1/10 Explaining In-Context Learning as Bayesian Updating Martin notes that in-context learning was non-obvious to him and shares his experience using Vishal's domain-specific language (DSL). Vishal explains how showing few-shot examples continuously increases target token probabilities, demonstrating real-time Bayesian updating.15:02–23:02 · Guest disagreement 2/10 Proving Bayesian Behavior via the Bayesian Wind Tunnel Martin presses on why a second paper was necessary when the first paper already seemed empirically convincing. Vishal explains the community skepticism stemming from Bayesian versus frequentist debates and outlines how the Bayesian wind tunnel mathematically proved exact Bayesian posterior matching.23:02–30:19 · Guest disagreement 5/10 Human Cognition vs. LLM Mechanics and Consciousness Martin brings up public claims about potential LLM consciousness. Vishal forcefully rejects the premise, emphasizing that LLMs are frozen matrix multiplications driven by next-token optimization rather than plastic human brains driven by survival.30:19–38:54 · Guest disagreement 2/10 Why Scale Will Not Solve AGI and The Einstein Test Martin demonstrates strong domain insight by articulating how LLMs suffer from 'data gravity,' binding them to existing training consensus. Vishal uses Einstein's relativity theory to show why scaling existing correlation models cannot produce new Kolmogorov representations.38:54–46:25 · Guest disagreement 1/10 Causal Models, Knuth's Experiment, and Future Research Martin queries whether simulation mechanisms align with Kolmogorov complexity and asks about practical research directions. Vishal breaks down Donald Knuth's viral LLM experiment, showing that Knuth manually supplied the causal reasoning while the LLM performed association.0:42–2:55 · The host pushing back 1/10 Early AI Experiments and Origin of RAG Martin opens by welcoming back Vishal and framing his work on LLM mathematical modeling as top-tier. Vishal shares the origin story of using GPT-3 in 2020 to build an early RAG implementation for ESPN cricket stats.2:55–8:36 · The host pushing back 1/10 The Matrix Abstraction of Language Models Martin asks targeted questions about row combinations and posterior distributions within Vishal's matrix model abstraction. Vishal explains how prompt rows generate vocabulary probability distributions and how huge context windows create astronomically large sparse matrices.8:36–15:02 · The host pushing back 2/10 Explaining In-Context Learning as Bayesian Updating Martin notes that in-context learning was non-obvious to him and shares his experience using Vishal's domain-specific language (DSL). Vishal explains how showing few-shot examples continuously increases target token probabilities, demonstrating real-time Bayesian updating.15:02–23:02 · The host pushing back 3/10 Proving Bayesian Behavior via the Bayesian Wind Tunnel Martin presses on why a second paper was necessary when the first paper already seemed empirically convincing. Vishal explains the community skepticism stemming from Bayesian versus frequentist debates and outlines how the Bayesian wind tunnel mathematically proved exact Bayesian posterior matching.23:02–30:19 · The host pushing back 2/10 Human Cognition vs. LLM Mechanics and Consciousness Martin brings up public claims about potential LLM consciousness. Vishal forcefully rejects the premise, emphasizing that LLMs are frozen matrix multiplications driven by next-token optimization rather than plastic human brains driven by survival.30:19–38:54 · The host pushing back 2/10 Why Scale Will Not Solve AGI and The Einstein Test Martin demonstrates strong domain insight by articulating how LLMs suffer from 'data gravity,' binding them to existing training consensus. Vishal uses Einstein's relativity theory to show why scaling existing correlation models cannot produce new Kolmogorov representations.38:54–46:25 · The host pushing back 2/10 Causal Models, Knuth's Experiment, and Future Research Martin queries whether simulation mechanisms align with Kolmogorov complexity and asks about practical research directions. Vishal breaks down Donald Knuth's viral LLM experiment, showing that Knuth manually supplied the causal reasoning while the LLM performed association.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 0% · guest 100%0:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%33:00 · the host 0% · guest 100%33:00 · the host 0% · guest 100%36:00 · the host 0% · guest 100%36:00 · the host 0% · guest 100%39:00 · the host 0% · guest 100%39:00 · the host 0% · guest 100%42:00 · the host 0% · guest 100%42:00 · the host 0% · guest 100%45:00 · the host 0% · guest 100%45:00 · the host 0% · guest 100%
Sharpest disagreement ▶ 26:05 Dismissing LLM Consciousness Claims

Vishal strongly rejects Dario Amodei's statement about LLM consciousness, exclaiming 'come on' and describing models as merely grains of silicon doing matrix multiplication.

Hardest push from the host ▶ 18:48 Questioning Need for Mathematical Proof

Martin challenges the necessity of the follow-up work, stating that the first paper already seemed conclusive and asking what was missing to warrant deeper proof.

Biggest teaching moment ▶ 23:18 Human Plasticity vs Frozen Weights

Vishal educates Martin on the structural divergence between human brain plasticity shaped by survival objectives and LLM static weights frozen after training.

The host holds their own ▶ 35:25 Formulating Data Gravity Concept

Martin demonstrates deep expertise by synthesizing Vishal's arguments into a concise theory of 'data gravity,' explaining why LLMs remain tethered to majority training data.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
Early AI Experiments and Origin of RAG 2311 Martin opens by welcoming back Vishal and framing his work on LLM mathematical modeling as top-tier. Vishal shares the origin story of using GPT-3 in 2020 to build an early RAG implementation for ESPN cricket stats.
The Matrix Abstraction of Language Models 3511 Martin asks targeted questions about row combinations and posterior distributions within Vishal's matrix model abstraction. Vishal explains how prompt rows generate vocabulary probability distributions and how huge context windows create astronomically large sparse matrices.
Explaining In-Context Learning as Bayesian Updating 3612 Martin notes that in-context learning was non-obvious to him and shares his experience using Vishal's domain-specific language (DSL). Vishal explains how showing few-shot examples continuously increases target token probabilities, demonstrating real-time Bayesian updating.
Proving Bayesian Behavior via the Bayesian Wind Tunnel 4623 Martin presses on why a second paper was necessary when the first paper already seemed empirically convincing. Vishal explains the community skepticism stemming from Bayesian versus frequentist debates and outlines how the Bayesian wind tunnel mathematically proved exact Bayesian posterior matching.
Human Cognition vs. LLM Mechanics and Consciousness 3652 Martin brings up public claims about potential LLM consciousness. Vishal forcefully rejects the premise, emphasizing that LLMs are frozen matrix multiplications driven by next-token optimization rather than plastic human brains driven by survival.
Why Scale Will Not Solve AGI and The Einstein Test 5522 Martin demonstrates strong domain insight by articulating how LLMs suffer from 'data gravity,' binding them to existing training consensus. Vishal uses Einstein's relativity theory to show why scaling existing correlation models cannot produce new Kolmogorov representations.
Causal Models, Knuth's Experiment, and Future Research 4512 Martin queries whether simulation mechanisms align with Kolmogorov complexity and asks about practical research directions. Vishal breaks down Donald Knuth's viral LLM experiment, showing that Knuth manually supplied the causal reasoning while the LLM performed association.

Statements from this episode (14)

Assertion Not checkable as stated
Misra claims he built the first known implementation of RAG using GPT-3
“And I got GPD three to do in context learning, few short learning, and you know, it was kind of the First, at least to me, it was the first known implementation of RAG, Retrieval Augmented Generation, which I used to solve this problem of querying, getting GPT…”
Vishal Misra Mar 17, 2026 ▶ 1:20
Assertion Not checkable as stated
ESPN deployed a GPT-3 database query system in production in September 2021
“We deployed this In production at ESPN in September, 21.”
Vishal Misra Mar 17, 2026 ▶ 1:55
Assertion Supported
LLM prompt combinations exceed the number of electrons in the universe
“If you look at all possible combinations of 8000 tokens and 50,000, ah, vocabulary, the number of rows In this matrix is more than the number of electrons across all galaxies, right? So, so there's no way that these LLMs can represent it exactly.”
Vishal Misra Mar 17, 2026 ▶ 6:35
Insight
Misra: LLMs function as compressed representations of prompt-probability matrices
“So in kind of an abstract way, what all these LLMs are doing is coming, coming up with a compressed representation of this matrix. And when you give a prompt, They try to approximate what the true distribution should have been and try to generate it.”
Vishal Misra Mar 17, 2026 ▶ 7:22
Assertion Supported
Transformers compute precise Bayesian posteriors down to 10^-3 bits accuracy
“We trained these models and we found that the transformer got the precise Bayesian posterior down to 10 to the power minus three bits accuracy. It was matching the distribution perfectly. So it is actually doing Bayesian in the mathematical sense, given a task…”
Vishal Misra Mar 17, 2026 ▶ 20:44
Assertion Supported
Transformers perform all Bayesian tasks, Mamba does most, and MLPs fail
“Transformer does everything. Mamba does most of it. LSTMs do only partially, and MLPs fail completely.”
Vishal Misra Mar 17, 2026 ▶ 21:13
Assertion Supported
Misra: Geometric Bayesian signatures persist in large open-weight LLMs
“We took these frontier production LLMs, which have open weights so that we could look inside them. And we did our testing and we saw that the geometries that we saw in the small models persisted in models, which are, you know, hundreds of millions of parameter…”
Vishal Misra Mar 17, 2026 ▶ 22:01
Assertion Not checkable as stated
Misra: LLM deception is driven by training data, not architecture
“That's not a function of the architecture. That's a function of the training data. It has been fed, you know, articles on Reddit or Asimo or whatever.”
Vishal Misra Mar 17, 2026 ▶ 25:47
Assertion Not checkable as stated
Misra: LLMs are just silicon doing matrix multiplication, not conscious beings
“You can rule out they're conscious. I mean, come on. And I said, you know, Anthropic makes great products. Cloud code is fantastic. Co-work is fantastic, but they are grains of silicon doing matrix multiplication. They don't have consciousness. They don't have…”
Vishal Misra Mar 17, 2026 ▶ 26:07
Insight
Misra: Deep learning performs correlation rather than causation
“All of deep learning is, ah, doing correlations. It's not doing causation. Causal models are the ones that are able to do simulations and intervention.”
Vishal Misra Mar 17, 2026 ▶ 28:03
Insight
Deep learning remains in the Shannon entropy world, lacking Kolmogorov complexity
“I think deep learning is still in the Shannon entropy world. It has not crossed over to the Kolmogorov complexity and the causal world.”
Vishal Misra Mar 17, 2026 ▶ 30:07
Prediction Not checkable as stated
Misra: Scale will not solve everything; AGI requires a different architecture
“One of the misconceptions that exists today is that scale will solve everything. Scale will not solve everything. You need a different kind of architecture, and this continual learning is a difficult problem.”
Vishal Misra Mar 17, 2026 ▶ 31:09
Insight
Misra: AGI requires real-time plasticity and causal world models
“To get to what is called AGI, I think there are two things that need to happen. One is this plasticity. Which has to be implemented through container learning. Secondly, we have to move from correlation to causation.”
Vishal Misra Mar 17, 2026 ▶ 31:46
Prediction Open · timeframe Mar 2031
An LLM trained on pre-1916 physics will never derive general relativity
“Take an LLM and train it on pre-nineteen-sixteen or nineteen-eleven physics and see if it can come up with the theory of relativity. If it does, then we have AGI. I mean, it's a high bar, but, you know, we should have high bars. It won't.”
Vishal Misra Mar 17, 2026 ▶ 32:31
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.