Dec 31, 2025 · 28m · latent-space

[NeurIPS Best Paper] 1000 Layer Networks for Self-Supervised RL — Kevin Wang et al, Princeton

Kevin Wang · 11m spoken Ishan Gaur · 3m spoken Benjamin Eysenbach · 2m spoken Michał Zawalski · 1m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Researchers Kevin Wang, Ishan Gaur, Michał Zawalski, and Professor Benjamin Eysenbach discuss their NeurIPS Best Paper award-winning research that successfully trains 1000-layer networks in reinforcement learning. They detail how replacing reward-based TD learning with self-supervised objectives and residual architectures enables unprecedented depth, computational efficiency, and scalability for embodied robotics and AI foundation models.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 4.5 Guest teaching 5.2 Guest disagreement 1.7 The hosts pushing back 2.0
05100:0010:0020:000:12–3:29 · The hosts as informed peer 3/10 Podcast Welcome and Celebrating the NeurIPS Best Paper Award The host warmly congratulates the authors on their NeurIPS Best Paper award and asks Ben Eysenbach how he chooses research bets. Ben explains the historical skepticism in deep RL where networks rarely exceeded three or four layers.3:30–7:40 · The hosts as informed peer 4/10 Scaling Self-Supervised RL via Depth and Residual Connections Kevin and Ishan explain how combining self-supervised objectives with residual connections and layer normalization unlocked scaling to 1000 layers. Ishan details how depth scaling maintains linear parameter growth compared to quadratic growth from width.7:40–13:10 · The hosts as informed peer 4/10 Reframing RL Objectives and Blurring Learning Paradigms Ben clarifies that the title is somewhat of a misnomer because the paper removes reward maximization entirely in favor of contrastive self-supervised learning. Kevin breaks down how shifting from temporal difference errors to classification enables massive scale.13:11–17:22 · The hosts as informed peer 3/10 Parameter Efficiency and Infrastructure Acceleration with JAX-GCRL The host asks whether depth scaling causes quadratic slowdowns, prompting Ishan to clarify the parameter growth formulas. Kevin and Michal explain how JAX-GCRL removes data collection bottlenecks by generating parallel rollouts on GPUs.17:23–22:04 · The hosts as informed peer 7/10 Analogies to Language Modeling, Implicit World Models, and Distillation The host actively debates LLM pre-training paradigms and proposes implicit world modeling analogies and a 'deep teacher, shallow student' distillation scheme. Kevin and Ishan validate these parallels, noting distillation is already on their future roadmap.22:04–26:43 · The hosts as informed peer 6/10 Future Research Horizons: Sub-Behavior Stitching, Multi-Axis Scaling, and VLAs The authors discuss sub-behavior stitching, multi-axis scaling on a single H100 GPU, and vision-language-action models. The host demonstrates domain knowledge by referencing recent industry interviews and discussing hierarchical action planning.0:12–3:29 · Guest teaching 3/10 Podcast Welcome and Celebrating the NeurIPS Best Paper Award The host warmly congratulates the authors on their NeurIPS Best Paper award and asks Ben Eysenbach how he chooses research bets. Ben explains the historical skepticism in deep RL where networks rarely exceeded three or four layers.3:30–7:40 · Guest teaching 6/10 Scaling Self-Supervised RL via Depth and Residual Connections Kevin and Ishan explain how combining self-supervised objectives with residual connections and layer normalization unlocked scaling to 1000 layers. Ishan details how depth scaling maintains linear parameter growth compared to quadratic growth from width.7:40–13:10 · Guest teaching 6/10 Reframing RL Objectives and Blurring Learning Paradigms Ben clarifies that the title is somewhat of a misnomer because the paper removes reward maximization entirely in favor of contrastive self-supervised learning. Kevin breaks down how shifting from temporal difference errors to classification enables massive scale.13:11–17:22 · Guest teaching 6/10 Parameter Efficiency and Infrastructure Acceleration with JAX-GCRL The host asks whether depth scaling causes quadratic slowdowns, prompting Ishan to clarify the parameter growth formulas. Kevin and Michal explain how JAX-GCRL removes data collection bottlenecks by generating parallel rollouts on GPUs.17:23–22:04 · Guest teaching 5/10 Analogies to Language Modeling, Implicit World Models, and Distillation The host actively debates LLM pre-training paradigms and proposes implicit world modeling analogies and a 'deep teacher, shallow student' distillation scheme. Kevin and Ishan validate these parallels, noting distillation is already on their future roadmap.22:04–26:43 · Guest teaching 5/10 Future Research Horizons: Sub-Behavior Stitching, Multi-Axis Scaling, and VLAs The authors discuss sub-behavior stitching, multi-axis scaling on a single H100 GPU, and vision-language-action models. The host demonstrates domain knowledge by referencing recent industry interviews and discussing hierarchical action planning.0:12–3:29 · Guest disagreement 1/10 Podcast Welcome and Celebrating the NeurIPS Best Paper Award The host warmly congratulates the authors on their NeurIPS Best Paper award and asks Ben Eysenbach how he chooses research bets. Ben explains the historical skepticism in deep RL where networks rarely exceeded three or four layers.3:30–7:40 · Guest disagreement 1/10 Scaling Self-Supervised RL via Depth and Residual Connections Kevin and Ishan explain how combining self-supervised objectives with residual connections and layer normalization unlocked scaling to 1000 layers. Ishan details how depth scaling maintains linear parameter growth compared to quadratic growth from width.7:40–13:10 · Guest disagreement 2/10 Reframing RL Objectives and Blurring Learning Paradigms Ben clarifies that the title is somewhat of a misnomer because the paper removes reward maximization entirely in favor of contrastive self-supervised learning. Kevin breaks down how shifting from temporal difference errors to classification enables massive scale.13:11–17:22 · Guest disagreement 2/10 Parameter Efficiency and Infrastructure Acceleration with JAX-GCRL The host asks whether depth scaling causes quadratic slowdowns, prompting Ishan to clarify the parameter growth formulas. Kevin and Michal explain how JAX-GCRL removes data collection bottlenecks by generating parallel rollouts on GPUs.17:23–22:04 · Guest disagreement 3/10 Analogies to Language Modeling, Implicit World Models, and Distillation The host actively debates LLM pre-training paradigms and proposes implicit world modeling analogies and a 'deep teacher, shallow student' distillation scheme. Kevin and Ishan validate these parallels, noting distillation is already on their future roadmap.22:04–26:43 · Guest disagreement 1/10 Future Research Horizons: Sub-Behavior Stitching, Multi-Axis Scaling, and VLAs The authors discuss sub-behavior stitching, multi-axis scaling on a single H100 GPU, and vision-language-action models. The host demonstrates domain knowledge by referencing recent industry interviews and discussing hierarchical action planning.0:12–3:29 · The hosts pushing back 1/10 Podcast Welcome and Celebrating the NeurIPS Best Paper Award The host warmly congratulates the authors on their NeurIPS Best Paper award and asks Ben Eysenbach how he chooses research bets. Ben explains the historical skepticism in deep RL where networks rarely exceeded three or four layers.3:30–7:40 · The hosts pushing back 1/10 Scaling Self-Supervised RL via Depth and Residual Connections Kevin and Ishan explain how combining self-supervised objectives with residual connections and layer normalization unlocked scaling to 1000 layers. Ishan details how depth scaling maintains linear parameter growth compared to quadratic growth from width.7:40–13:10 · The hosts pushing back 2/10 Reframing RL Objectives and Blurring Learning Paradigms Ben clarifies that the title is somewhat of a misnomer because the paper removes reward maximization entirely in favor of contrastive self-supervised learning. Kevin breaks down how shifting from temporal difference errors to classification enables massive scale.13:11–17:22 · The hosts pushing back 2/10 Parameter Efficiency and Infrastructure Acceleration with JAX-GCRL The host asks whether depth scaling causes quadratic slowdowns, prompting Ishan to clarify the parameter growth formulas. Kevin and Michal explain how JAX-GCRL removes data collection bottlenecks by generating parallel rollouts on GPUs.17:23–22:04 · The hosts pushing back 4/10 Analogies to Language Modeling, Implicit World Models, and Distillation The host actively debates LLM pre-training paradigms and proposes implicit world modeling analogies and a 'deep teacher, shallow student' distillation scheme. Kevin and Ishan validate these parallels, noting distillation is already on their future roadmap.22:04–26:43 · The hosts pushing back 2/10 Future Research Horizons: Sub-Behavior Stitching, Multi-Axis Scaling, and VLAs The authors discuss sub-behavior stitching, multi-axis scaling on a single H100 GPU, and vision-language-action models. The host demonstrates domain knowledge by referencing recent industry interviews and discussing hierarchical action planning.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 18:20 Kevin correcting the host on changing LLM pretraining

Kevin directly refutes the host's premise that they intend to modify LLM pre-training, clarifying that they are borrowing insights from language modeling to improve RL.

Hardest push from the hosts ▶ 18:18 Host challenging language model analogies

The host questions the next-token analogy and pushes back, arguing that the flow of research insights should travel in the opposite direction.

Biggest teaching moment ▶ 7:53 Ben recontextualizing the entire paper's premise

Ben explains to the host and audience that scaling required throwing out conventional reward maximization entirely, calling the term 'RL' in their own title a misnomer.

The host holds their own ▶ 20:00 Host formulating implicit world modeling and distillation strategy

The host demonstrates strong conceptual mastery by framing goal-conditioned contrastive learning as implicit world modeling and proposing a teacher-student distillation setup matching the authors' roadmap.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Podcast Welcome and Celebrating the NeurIPS Best Paper Award 3311 The host warmly congratulates the authors on their NeurIPS Best Paper award and asks Ben Eysenbach how he chooses research bets. Ben explains the historical skepticism in deep RL where networks rarely exceeded three or four layers.
Scaling Self-Supervised RL via Depth and Residual Connections 4611 Kevin and Ishan explain how combining self-supervised objectives with residual connections and layer normalization unlocked scaling to 1000 layers. Ishan details how depth scaling maintains linear parameter growth compared to quadratic growth from width.
Reframing RL Objectives and Blurring Learning Paradigms 4622 Ben clarifies that the title is somewhat of a misnomer because the paper removes reward maximization entirely in favor of contrastive self-supervised learning. Kevin breaks down how shifting from temporal difference errors to classification enables massive scale.
Parameter Efficiency and Infrastructure Acceleration with JAX-GCRL 3622 The host asks whether depth scaling causes quadratic slowdowns, prompting Ishan to clarify the parameter growth formulas. Kevin and Michal explain how JAX-GCRL removes data collection bottlenecks by generating parallel rollouts on GPUs.
Analogies to Language Modeling, Implicit World Models, and Distillation 7534 The host actively debates LLM pre-training paradigms and proposes implicit world modeling analogies and a 'deep teacher, shallow student' distillation scheme. Kevin and Ishan validate these parallels, noting distillation is already on their future roadmap.
Future Research Horizons: Sub-Behavior Stitching, Multi-Axis Scaling, and VLAs 6512 The authors discuss sub-behavior stitching, multi-axis scaling on a single H100 GPU, and vision-language-action models. The host demonstrates domain knowledge by referencing recent industry interviews and discussing hierarchical action planning.

Statements from this episode (12)

Assertion Supported
Eysenbach: Historically, 'deep' RL meant only two to four layers
“So it's like, probably my lab works on deep reinforcement learning, but historically deep meant like two or three or four layers.”
Benjamin Eysenbach Dec 31, 2025 ▶ 2:15
Assertion Not checkable as stated
Wang: Traditional value-based reinforcement learning fails to scale
“And so what we did is that we know that traditional RL, like let's say like value value-based RL doesn't really scale, right? This is pretty clear from the literature.”
Kevin Wang Dec 31, 2025 ▶ 4:29
Assertion Supported
Eysenbach: Scaling RL depth requires combining depth with residual connections
“And if we just made the depth bigger, it makes it worse. If we just add residual connections, it didn't make it better. And it was really this combination of factors that Kevin and Ishan figured out that really made this work.”
Benjamin Eysenbach Dec 31, 2025 ▶ 5:56
Assertion Supported
Gaur: Scaling depth is more parameter- and sample-efficient than width in RL
“But when you look at the number of parameters that your network has as you grow with, it's roughly a quadratic as opposed to something like growing depth, so it's more, in some sense, it's more parameter efficient, also more sample efficient from the experimen…”
Ishan Gaur Dec 31, 2025 ▶ 6:42
Insight
Eysenbach: 1,000-layer RL requires reward-free objectives, not just architectural tricks
“I think the main conclusion is that using big networks not only requires these architectural tricks, but also, as Kevin mentioned before, it requires using a different objective. This objective doesn't actually use rewards in it, and so there's another word in…”
Benjamin Eysenbach Dec 31, 2025 ▶ 8:08
Insight
Kevin Wang: Cross-entropy trajectory classification enables scalable deep reinforcement learning
“I think it's because we're fundamentally shifting the burden of learning from something like, Q-learning or, like, regressing to, like, TD errors, which we know is quite spurious and noisy and biased, to fundamentally, like, a classification problem. We're try…”
Kevin Wang Dec 31, 2025 ▶ 9:56
Assertion Supported
Wang: 64 layers saturate performance in most reinforcement learning tasks
“Within our paper, like, for most environments we are able to, like, saturate, like, get to, like, almost perfect performance within just, you know, we don't even need to get to, like, a thousand layers. Like, maybe just 64 layers, for example, is sufficient.”
Kevin Wang Dec 31, 2025 ▶ 15:25
Assertion Partly supported
Michał Zawalski: RL Scaling Gains Require Over 50 Million Transitions
“Going back to our paper, if you look at the plots, we only see this, like, huge performance increase When we cross, like, 50 millions of transitions gap.”
Michał Zawalski Dec 31, 2025 ▶ 17:06
Assertion Supported
Wang: GPU environments collect hundreds of millions of RL timesteps hourly
“With these, like, GPU accelerated environments, we can collect hundreds of millions of time steps of data within just a few hours”
Kevin Wang Dec 31, 2025 ▶ 17:43
Assertion Supported
Wang: Deep Self-Supervised Method Beats SOTA on Goal-Conditioned RL Significantly
“We do achieve state-of-the-art performance on goal-conditioned RL and Jack's CCRL by a significant amount.”
Kevin Wang Dec 31, 2025 ▶ 21:24
Assertion Supported
Kevin Wang: Scaling RL network depth unlocks effective batch size scaling
“We notice that we see that scaling width actually also improves performance, and we also find that actually by scaling depth, we actually unlock the ability to scale along batch size as well.”
Kevin Wang Dec 31, 2025 ▶ 22:38
Assertion Supported
Kevin Wang: 1000-layer RL networks can train on single 80GB H100
“The nice thing is that all of our experiments, even the thousand layer networks, can be run on one single, 80 gigabyte, each 100 GPU.”
Kevin Wang Dec 31, 2025 ▶ 24:19
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.