May 1, 2026 · 37m · y-combinator

Recursion Is The Next Scaling Law In AI · Y Combinator

Francois Chaubard · 23m spoken Ankit Gupta · 11m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of YC Decoded, hosts Ankit Gupta and Francois Chaubard explore how test-time recursive compute depth in continuous latent space—highlighted by Hierarchical Reasoning Models (HRM) and Tiny Recursive Models (TRM)—offers a powerful alternative to traditional parameter scaling for complex AI reasoning.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The partners as informed peer 5.9 Guest teaching 3.5 Guest disagreement 1.1 The partners pushing back 1.4
05100:0010:0020:0030:000:35–3:46 · The partners as informed peer 6/10 Historical Evolution of Recurrent Neural Networks Ankit Gupta demonstrates a solid technical foundation in RNN mechanics, discussing vanishing gradients, matrix multiplications across time steps, and parallelization via causal masking. Francois builds collaboratively on this foundation, providing historical context around BPTT and memory trade-offs.3:46–6:20 · The partners as informed peer 6/10 One-Shot LLM Limitations on Incompressible Problems Gupta connects Francois's theoretical sorting comparison lower bounds to computer science fundamentals like external memory caches and radix sort. The dynamic is highly collaborative, with both speakers riffing on computational complexity.6:20–8:38 · The partners as informed peer 5/10 Turing Completeness, Chain of Thought, and Training Traces Gupta introduces the Turing completeness analogy while Francois clarifies how test-time chain of thought acts as an expressive hack bounded by human training traces.8:38–14:17 · The partners as informed peer 6/10 HRM Mechanics, Recursion Loops, and ArcPrize Breakthroughs Gupta tracks the three nested recursion loops in HRM and proactively brings up the BPTT bottleneck, prompting Francois to introduce deep equilibrium fixed-point iteration.14:17–16:45 · The partners as informed peer 7/10 HRM's Fixed-Point Iteration (DEQ) Trick for BPTT Gupta takes a clear stance pushing back against bio-plausibility arguments in deep learning, citing historical precedents of biologically implausible architectures winning on GPUs. Francois yields and asks the host for his expert opinion on the matter.16:45–19:00 · The partners as informed peer 6/10 Automata Theory and Memory Caching in Latent Space Gupta articulates his framework linking latent memory states to automata theory, while Francois points out that chain of thought cannot invent novel algorithms without prior demonstrations.19:00–22:05 · The partners as informed peer 6/10 Discrete Token Space vs. Continuous Latent Space Recursion Gupta clearly explains the key limitation of modern LLMs recursing strictly in discrete token space rather than continuous latent space, which Francois affirms and expands upon.22:05–27:39 · The partners as informed peer 6/10 TRM Architecture, Truncated BPTT, and Latent EM Optimization Gupta synthesizes the parameter depth versus compute depth trade-off, framing the latent state optimization as expectation-maximization. Francois validates the analogy using Sudoku as a concrete problem.27:39–30:36 · The partners as informed peer 5/10 Code Analysis: HRM PyTorch Implementation Walkthrough The host and guest walk through the PyTorch implementation of HRM, with Gupta actively pointing out how gradient detachment effectively creates mini-batches across memory carry space.30:36–34:39 · The partners as informed peer 6/10 Code Analysis: TRM Implementation and Parameter Scaling Gupta observes that TRM simplifies HRM by weight sharing and reducing layer count while expanding BPTT depth. Francois emphasizes parameter efficiency on ArcPrize benchmarks.34:39–37:44 · The partners as informed peer 6/10 The Future of AI: Merging Giant LLMs with Latent Recursion Gupta concludes with a nuanced distinction between task-specific recursive models and general-purpose foundation models, speculating on their eventual architectural convergence.0:35–3:46 · Guest teaching 3/10 Historical Evolution of Recurrent Neural Networks Ankit Gupta demonstrates a solid technical foundation in RNN mechanics, discussing vanishing gradients, matrix multiplications across time steps, and parallelization via causal masking. Francois builds collaboratively on this foundation, providing historical context around BPTT and memory trade-offs.3:46–6:20 · Guest teaching 3/10 One-Shot LLM Limitations on Incompressible Problems Gupta connects Francois's theoretical sorting comparison lower bounds to computer science fundamentals like external memory caches and radix sort. The dynamic is highly collaborative, with both speakers riffing on computational complexity.6:20–8:38 · Guest teaching 4/10 Turing Completeness, Chain of Thought, and Training Traces Gupta introduces the Turing completeness analogy while Francois clarifies how test-time chain of thought acts as an expressive hack bounded by human training traces.8:38–14:17 · Guest teaching 5/10 HRM Mechanics, Recursion Loops, and ArcPrize Breakthroughs Gupta tracks the three nested recursion loops in HRM and proactively brings up the BPTT bottleneck, prompting Francois to introduce deep equilibrium fixed-point iteration.14:17–16:45 · Guest teaching 2/10 HRM's Fixed-Point Iteration (DEQ) Trick for BPTT Gupta takes a clear stance pushing back against bio-plausibility arguments in deep learning, citing historical precedents of biologically implausible architectures winning on GPUs. Francois yields and asks the host for his expert opinion on the matter.16:45–19:00 · Guest teaching 4/10 Automata Theory and Memory Caching in Latent Space Gupta articulates his framework linking latent memory states to automata theory, while Francois points out that chain of thought cannot invent novel algorithms without prior demonstrations.19:00–22:05 · Guest teaching 3/10 Discrete Token Space vs. Continuous Latent Space Recursion Gupta clearly explains the key limitation of modern LLMs recursing strictly in discrete token space rather than continuous latent space, which Francois affirms and expands upon.22:05–27:39 · Guest teaching 4/10 TRM Architecture, Truncated BPTT, and Latent EM Optimization Gupta synthesizes the parameter depth versus compute depth trade-off, framing the latent state optimization as expectation-maximization. Francois validates the analogy using Sudoku as a concrete problem.27:39–30:36 · Guest teaching 4/10 Code Analysis: HRM PyTorch Implementation Walkthrough The host and guest walk through the PyTorch implementation of HRM, with Gupta actively pointing out how gradient detachment effectively creates mini-batches across memory carry space.30:36–34:39 · Guest teaching 4/10 Code Analysis: TRM Implementation and Parameter Scaling Gupta observes that TRM simplifies HRM by weight sharing and reducing layer count while expanding BPTT depth. Francois emphasizes parameter efficiency on ArcPrize benchmarks.34:39–37:44 · Guest teaching 3/10 The Future of AI: Merging Giant LLMs with Latent Recursion Gupta concludes with a nuanced distinction between task-specific recursive models and general-purpose foundation models, speculating on their eventual architectural convergence.0:35–3:46 · Guest disagreement 1/10 Historical Evolution of Recurrent Neural Networks Ankit Gupta demonstrates a solid technical foundation in RNN mechanics, discussing vanishing gradients, matrix multiplications across time steps, and parallelization via causal masking. Francois builds collaboratively on this foundation, providing historical context around BPTT and memory trade-offs.3:46–6:20 · Guest disagreement 1/10 One-Shot LLM Limitations on Incompressible Problems Gupta connects Francois's theoretical sorting comparison lower bounds to computer science fundamentals like external memory caches and radix sort. The dynamic is highly collaborative, with both speakers riffing on computational complexity.6:20–8:38 · Guest disagreement 1/10 Turing Completeness, Chain of Thought, and Training Traces Gupta introduces the Turing completeness analogy while Francois clarifies how test-time chain of thought acts as an expressive hack bounded by human training traces.8:38–14:17 · Guest disagreement 1/10 HRM Mechanics, Recursion Loops, and ArcPrize Breakthroughs Gupta tracks the three nested recursion loops in HRM and proactively brings up the BPTT bottleneck, prompting Francois to introduce deep equilibrium fixed-point iteration.14:17–16:45 · Guest disagreement 2/10 HRM's Fixed-Point Iteration (DEQ) Trick for BPTT Gupta takes a clear stance pushing back against bio-plausibility arguments in deep learning, citing historical precedents of biologically implausible architectures winning on GPUs. Francois yields and asks the host for his expert opinion on the matter.16:45–19:00 · Guest disagreement 1/10 Automata Theory and Memory Caching in Latent Space Gupta articulates his framework linking latent memory states to automata theory, while Francois points out that chain of thought cannot invent novel algorithms without prior demonstrations.19:00–22:05 · Guest disagreement 1/10 Discrete Token Space vs. Continuous Latent Space Recursion Gupta clearly explains the key limitation of modern LLMs recursing strictly in discrete token space rather than continuous latent space, which Francois affirms and expands upon.22:05–27:39 · Guest disagreement 1/10 TRM Architecture, Truncated BPTT, and Latent EM Optimization Gupta synthesizes the parameter depth versus compute depth trade-off, framing the latent state optimization as expectation-maximization. Francois validates the analogy using Sudoku as a concrete problem.27:39–30:36 · Guest disagreement 1/10 Code Analysis: HRM PyTorch Implementation Walkthrough The host and guest walk through the PyTorch implementation of HRM, with Gupta actively pointing out how gradient detachment effectively creates mini-batches across memory carry space.30:36–34:39 · Guest disagreement 1/10 Code Analysis: TRM Implementation and Parameter Scaling Gupta observes that TRM simplifies HRM by weight sharing and reducing layer count while expanding BPTT depth. Francois emphasizes parameter efficiency on ArcPrize benchmarks.34:39–37:44 · Guest disagreement 1/10 The Future of AI: Merging Giant LLMs with Latent Recursion Gupta concludes with a nuanced distinction between task-specific recursive models and general-purpose foundation models, speculating on their eventual architectural convergence.0:35–3:46 · The partners pushing back 1/10 Historical Evolution of Recurrent Neural Networks Ankit Gupta demonstrates a solid technical foundation in RNN mechanics, discussing vanishing gradients, matrix multiplications across time steps, and parallelization via causal masking. Francois builds collaboratively on this foundation, providing historical context around BPTT and memory trade-offs.3:46–6:20 · The partners pushing back 2/10 One-Shot LLM Limitations on Incompressible Problems Gupta connects Francois's theoretical sorting comparison lower bounds to computer science fundamentals like external memory caches and radix sort. The dynamic is highly collaborative, with both speakers riffing on computational complexity.6:20–8:38 · The partners pushing back 1/10 Turing Completeness, Chain of Thought, and Training Traces Gupta introduces the Turing completeness analogy while Francois clarifies how test-time chain of thought acts as an expressive hack bounded by human training traces.8:38–14:17 · The partners pushing back 2/10 HRM Mechanics, Recursion Loops, and ArcPrize Breakthroughs Gupta tracks the three nested recursion loops in HRM and proactively brings up the BPTT bottleneck, prompting Francois to introduce deep equilibrium fixed-point iteration.14:17–16:45 · The partners pushing back 3/10 HRM's Fixed-Point Iteration (DEQ) Trick for BPTT Gupta takes a clear stance pushing back against bio-plausibility arguments in deep learning, citing historical precedents of biologically implausible architectures winning on GPUs. Francois yields and asks the host for his expert opinion on the matter.16:45–19:00 · The partners pushing back 1/10 Automata Theory and Memory Caching in Latent Space Gupta articulates his framework linking latent memory states to automata theory, while Francois points out that chain of thought cannot invent novel algorithms without prior demonstrations.19:00–22:05 · The partners pushing back 1/10 Discrete Token Space vs. Continuous Latent Space Recursion Gupta clearly explains the key limitation of modern LLMs recursing strictly in discrete token space rather than continuous latent space, which Francois affirms and expands upon.22:05–27:39 · The partners pushing back 1/10 TRM Architecture, Truncated BPTT, and Latent EM Optimization Gupta synthesizes the parameter depth versus compute depth trade-off, framing the latent state optimization as expectation-maximization. Francois validates the analogy using Sudoku as a concrete problem.27:39–30:36 · The partners pushing back 1/10 Code Analysis: HRM PyTorch Implementation Walkthrough The host and guest walk through the PyTorch implementation of HRM, with Gupta actively pointing out how gradient detachment effectively creates mini-batches across memory carry space.30:36–34:39 · The partners pushing back 1/10 Code Analysis: TRM Implementation and Parameter Scaling Gupta observes that TRM simplifies HRM by weight sharing and reducing layer count while expanding BPTT depth. Francois emphasizes parameter efficiency on ArcPrize benchmarks.34:39–37:44 · The partners pushing back 1/10 The Future of AI: Merging Giant LLMs with Latent Recursion Gupta concludes with a nuanced distinction between task-specific recursive models and general-purpose foundation models, speculating on their eventual architectural convergence.

speaking balance: gold is the partners, purple is the guest (3 minute bins)

0:00 · the partners 0% · guest 100%0:00 · the partners 0% · guest 100%3:00 · the partners 0% · guest 100%3:00 · the partners 0% · guest 100%6:00 · the partners 0% · guest 100%6:00 · the partners 0% · guest 100%9:00 · the partners 0% · guest 100%9:00 · the partners 0% · guest 100%12:00 · the partners 0% · guest 100%12:00 · the partners 0% · guest 100%15:00 · the partners 0% · guest 100%15:00 · the partners 0% · guest 100%18:00 · the partners 0% · guest 100%18:00 · the partners 0% · guest 100%21:00 · the partners 0% · guest 100%21:00 · the partners 0% · guest 100%24:00 · the partners 0% · guest 100%24:00 · the partners 0% · guest 100%27:00 · the partners 0% · guest 100%27:00 · the partners 0% · guest 100%30:00 · the partners 0% · guest 100%30:00 · the partners 0% · guest 100%33:00 · the partners 0% · guest 100%33:00 · the partners 0% · guest 100%36:00 · the partners 0% · guest 100%36:00 · the partners 0% · guest 100%
Sharpest disagreement ▶ 18:18 Guest dismisses standard CoT as bounded by human knowledge

Francois bluntly points out that if a problem requires knowledge outside human demonstration, standard chain-of-thought methods leave developers completely stuck.

Hardest push from the partners ▶ 14:20 Host rejects bio-plausibility as an architectural goal

Ankit directly challenges the utility of bio-plausible architectures, emphasizing that machine learning history proves biologically implausible GPU-optimized variants consistently perform better.

Biggest teaching moment ▶ 11:57 Guest explains DEQ fixed-point iteration trick

Francois educates the host on how HRM uses deep equilibrium pseudo fixed-point iteration to update weights without unfolding full BPTT across all recursion steps.

The partners hold their own ▶ 24:54 Host maps TRM latent updating to expectation-maximization

Ankit demonstrates deep technical mastery by independently translating TRM's dual-latent update cycle into an expectation-maximization framework.

the scores for every segment, with the reasoning behind each
ChapterTopicThe partners as informed peerGuest teachingGuest disagreementThe partners pushing backWhy
Historical Evolution of Recurrent Neural Networks 6311 Ankit Gupta demonstrates a solid technical foundation in RNN mechanics, discussing vanishing gradients, matrix multiplications across time steps, and parallelization via causal masking. Francois builds collaboratively on this foundation, providing historical context around BPTT and memory trade-offs.
One-Shot LLM Limitations on Incompressible Problems 6312 Gupta connects Francois's theoretical sorting comparison lower bounds to computer science fundamentals like external memory caches and radix sort. The dynamic is highly collaborative, with both speakers riffing on computational complexity.
Turing Completeness, Chain of Thought, and Training Traces 5411 Gupta introduces the Turing completeness analogy while Francois clarifies how test-time chain of thought acts as an expressive hack bounded by human training traces.
HRM Mechanics, Recursion Loops, and ArcPrize Breakthroughs 6512 Gupta tracks the three nested recursion loops in HRM and proactively brings up the BPTT bottleneck, prompting Francois to introduce deep equilibrium fixed-point iteration.
HRM's Fixed-Point Iteration (DEQ) Trick for BPTT 7223 Gupta takes a clear stance pushing back against bio-plausibility arguments in deep learning, citing historical precedents of biologically implausible architectures winning on GPUs. Francois yields and asks the host for his expert opinion on the matter.
Automata Theory and Memory Caching in Latent Space 6411 Gupta articulates his framework linking latent memory states to automata theory, while Francois points out that chain of thought cannot invent novel algorithms without prior demonstrations.
Discrete Token Space vs. Continuous Latent Space Recursion 6311 Gupta clearly explains the key limitation of modern LLMs recursing strictly in discrete token space rather than continuous latent space, which Francois affirms and expands upon.
TRM Architecture, Truncated BPTT, and Latent EM Optimization 6411 Gupta synthesizes the parameter depth versus compute depth trade-off, framing the latent state optimization as expectation-maximization. Francois validates the analogy using Sudoku as a concrete problem.
Code Analysis: HRM PyTorch Implementation Walkthrough 5411 The host and guest walk through the PyTorch implementation of HRM, with Gupta actively pointing out how gradient detachment effectively creates mini-batches across memory carry space.
Code Analysis: TRM Implementation and Parameter Scaling 6411 Gupta observes that TRM simplifies HRM by weight sharing and reducing layer count while expanding BPTT depth. Francois emphasizes parameter efficiency on ArcPrize benchmarks.
The Future of AI: Merging Giant LLMs with Latent Recursion 6311 Gupta concludes with a nuanced distinction between task-specific recursive models and general-purpose foundation models, speculating on their eventual architectural convergence.

Statements from this episode (19)

Assertion Not checkable as stated
Chaubard: Prior to 2016, Researchers Believed RNNs Were Necessary for AGI
“An RNN is just a model that you recursively call again and again and again on itself, and we, We're very much in the belief that this was required to get to AGI peak RNN use was probably until 2016 with Alex Graves NeurIPS keynote, which is just fantastic, and…”
Francois Chaubard May 1, 2026 ▶ 0:44
Insight
Chaubard: Standard LLMs sacrifice latent time compression unlike RNNs
“What you actually paid for that you have to give up is this latent reasoning thing and this compression in the time direction. There is no compression in LMs. Every single decode that I do, I still have to retain the entire, you know, Shakespeare novel just to…”
Francois Chaubard May 1, 2026 ▶ 3:23
Insight
Chaubard: Transformers cannot sort lists longer than layer count in one pass
“In a one-shot basis. It's like literally that we know a theoretical lower bound that for comparison sort, you can't do better than n log n steps. And if I have a list that's 31 characters or elements long, and my transformer is 30, I run out of steps to do com…”
Francois Chaubard May 1, 2026 ▶ 4:56
Assertion Partly supported
Chaubard: Sudoku, mazes, and rolling sums are incompressible reasoning problems for LLMs
“In HRM and TRM, they use Sudoku as an incompressible problem. Similarly, and so are mazes. Those are incompressible problems. Rolling sum, incompressible problem.”
Francois Chaubard May 1, 2026 ▶ 5:20
Assertion Supported
Chauvard: Chain of thought makes LLMs Turing complete at test time
“And so it's completely true that at test time, they are turn complete. And you can simulate all turn computable functions at test time.”
Francois Chaubard May 1, 2026 ▶ 7:01
Insight
Chaubard: Chain of thought fails on unsolved problems lacking human traces
“Unless you're training it on human labeled traces for which there's a lot of problems like the millennial prize problem. We don't have the trace for it.”
Francois Chaubard May 1, 2026 ▶ 7:12
Opinion
Chaubard: Hierarchical Reasoning Models Offer Little Novelty Over Standard RNNs
“The, this is directly in the lineage of RNNs. There's not that much novel from, like, the RNN standpoint at least in my opinion.”
Francois Chaubard May 1, 2026 ▶ 7:45
Assertion Partly supported
Chaubard: HRM scored 70% on ARC Prize 1 without pre-training
“There is no pre-training at all. This starts from, like, literally Tagula-Rasa weights, and it can outperform at that time, if we go back, you know, we had O-three, if you remember back, way back when. And it, O-three gets zero. Literally zero, and this got, l…”
Francois Chaubard May 1, 2026 ▶ 10:09
Insight
Gupta: Machine learning progresses by abandoning bio-plausibility for computational efficiency
“I think machine learning tends to have a long history of people starting with bio-plausible arguments, and then realizing that there's some variant of them that seems highly bio-implausible that actually works better.”
Ankit Gupta May 1, 2026 ▶ 14:35
Insight
Gupta: Recursive Latent Hidden States Function Like a Turing Machine Tape
“I kind of think of this set of hidden states or carry as akin to a Turing machine tape or akin to the radix sort, ah, memory bank, where you can basically train a model to use this memory cache in an intelligent way in a single forward pass so that you can get…”
Ankit Gupta May 1, 2026 ▶ 17:03
Insight
Chaubard: Chain of thought trained on bubble sort won't discover merge sort
“If you chain of thought it on all the bubble sort input and output, it will only do bubble sort. In fact, it won't even do bubble That well.”
Francois Chaubard May 1, 2026 ▶ 18:48
Insight
Gupta: Chain of Thought Operates in Token Space, Not Model Recursion
“There already exists some type of recursion that people are used to in LLMs, which is a chain of thought we mentioned earlier, but that is a recursion that's happening in the token space of the model's outputs, not inherent to the model itself. That's sort of …”
Ankit Gupta May 1, 2026 ▶ 19:01
Insight
Gupta: Recursive architectures achieve compute depth without parameter depth
“Recursion advantage now gives you a bunch of advantages over transformers where rather than having, you know, 500 or a thousand or a million or whatever transformer layers and having tons and tons of parameters, you get compute depth basically without this par…”
Ankit Gupta May 1, 2026 ▶ 24:32
Insight
Chaubard: Recursive models discover problem-solving strategies without human teacher forcing
“That's the most important part, is that if we had Sudoku, and we know how to solve Sudoku, because like we were just, you know, dumb homo sapiens that didn't know how to solve Sudoku, like it would just have solved it. And that's why it's cool, because it actu…”
Francois Chaubard May 1, 2026 ▶ 27:20
Assertion Supported
Chaubard: Recursive models tested on one step retain nearly full performance
“If you actually train on 16, and you test on only one, you get, like, seven eighths of the performance, or, like, almost all the performance. So it's actually quite interesting that this is just overdone, too much compute, and it doesn't actually help you all …”
Francois Chaubard May 1, 2026 ▶ 30:01
Assertion Supported
Chaubard: MLP outperformed attention on Sudoku but scored zero on mazes
“Yeah, on Sudoku, MLP actually outperformed the attention, It was it scored zero on the maze, ah, the MLP scored zero on the maze, and so there's, it's not clear, it's not obvious that, ah, the transformer is always better.”
Francois Chaubard May 1, 2026 ▶ 31:50
Assertion Contradicted
Chaubard: 7M parameter TRM scored 87% on ARC Prize 1
“And so it's a twenty-eight million parameter model for HRM. Now she brings it down to a seven million parameter model. It actually gets from 70% to 87% on on ArcPrize one. And does actually quite well on ArcPrize two as well.”
Francois Chaubard May 1, 2026 ▶ 33:25
Assertion Supported
Gupta: Current TRMs and HRMs are task-specific, not general-purpose
“One of the things that's really interesting about these TRMs and HRMs is they're not general purpose models, right? These were Task specific models, right? The model trained to do Sudoku cannot do ArcPrize inherently. It has to be trained on the ArcPrize set t…”
Ankit Gupta May 1, 2026 ▶ 36:19
Prediction Not checkable as stated
Chaubard predicts running recursive reasoning models inside giant LLM latent spaces
“What you can imagine is we found mapping from token space or from vision, from pixels, Some really cool latent space where, like, things are just nicely semantically separated, and we can, you know, makes it really easy for downstream tasks to do, but now in t…”
Francois Chaubard May 1, 2026 ▶ 37:12
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.