Mar 23, 2025 · 46m · latent-space

The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind

Rishabh Agarwal · 37m spoken Shawn Wang · 3m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Google DeepMind Staff Research Scientist Rishabh Agarwal joins the Latent Space podcast to break down the technical resurgence of LLM distillation, detailing how logit-based matching, on-policy generation, and reinforcement learning techniques optimize cost and performance for real-world deployments.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 4.9 Guest teaching 5.9 Guest disagreement 1.0 The hosts pushing back 2.0
05100:0015:0030:0045:005:48–8:51 · The hosts as informed peer 5/10 Dark Knowledge and Post-Training Pipeline Integration Alessio and Swyx explore the high-level motivation for distillation and ask whether post-training pipelines like RLHF can be bypassed. Rishabh explains how distillation fits into the post-training stack and defines Geoff Hinton's concept of dark knowledge.8:54–12:15 · The hosts as informed peer 4/10 Logit-Based Distillation and Soft Probability Labels Rishabh details classical logit-matching distillation and how soft probability labels supply richer supervision than hard one-hot tokens. Swyx raises a pertinent question regarding compute-matched versus data-matched training FLOPs.12:15–16:13 · The hosts as informed peer 4/10 Synthetic Data Distillation and Output Verification Rishabh outlines synthetic data distillation and how DeepSeek-R1 utilized best-of-n verifiable sampling across disparate tokenizers. Swyx asks about efficiency gaps between logit matching and sampling from black-box APIs.16:16–19:33 · The hosts as informed peer 5/10 Compute and Cost-Matched Data Generation Dynamics Rishabh shares empirical findings showing that generating filtered data from a smaller model like Gemma 9B beats self-generation from Gemma 27B under compute-matched conditions. Swyx succinctly captures the core intuition that filtering injects external signal.19:36–27:04 · The hosts as informed peer 5/10 Mitigating Exposure Bias via On-Policy Distillation Rishabh uses the DAgger driving analogy to explain exposure bias and train-test distribution shift when students make compounding errors. Swyx immediately identifies the phenomenon as exposure bias as Rishabh outlines the math of on-policy reverse KL distillation.27:06–31:35 · The hosts as informed peer 5/10 Theoretical Divergence: Mode-Covering Versus Mode-Seeking Rishabh presents the theoretical differences between mode-covering forward KL and mode-seeking reverse KL using Gaussian curves. Swyx pushes on the realism of toy Gaussian assumptions in natural language, prompting Rishabh to show real diversity-performance trade-off curves.31:38–34:04 · The hosts as informed peer 6/10 Repurposing RLHF Infrastructure for Model Distillation Rishabh explains how to implement on-policy distillation inside existing RLHF frameworks by zeroing the reward term and retargeting the KL anchor. Swyx directly challenges the framing, pressing Rishabh twice on why turning off rewards is considered profound.34:05–37:44 · The hosts as informed peer 4/10 Speculative Decoding Acceleration and Active Teacher Intervention Rishabh details how distillation accelerates speculative decoding and previews recent research on active teacher intervention during student rollout generation. Swyx synthesizes the takeaway that every large model deployment benefits from a distilled student companion.37:44–44:33 · The hosts as informed peer 6/10 Online Versus Offline Distillation for Complex Reasoning Swyx challenges the practical viability of online distillation, arguing that keeping a live teacher model during training defeats the cost-saving purpose. Rishabh counters firmly with the amortized serving economics of distilling once to serve billions of requests.5:48–8:51 · Guest teaching 5/10 Dark Knowledge and Post-Training Pipeline Integration Alessio and Swyx explore the high-level motivation for distillation and ask whether post-training pipelines like RLHF can be bypassed. Rishabh explains how distillation fits into the post-training stack and defines Geoff Hinton's concept of dark knowledge.8:54–12:15 · Guest teaching 6/10 Logit-Based Distillation and Soft Probability Labels Rishabh details classical logit-matching distillation and how soft probability labels supply richer supervision than hard one-hot tokens. Swyx raises a pertinent question regarding compute-matched versus data-matched training FLOPs.12:15–16:13 · Guest teaching 5/10 Synthetic Data Distillation and Output Verification Rishabh outlines synthetic data distillation and how DeepSeek-R1 utilized best-of-n verifiable sampling across disparate tokenizers. Swyx asks about efficiency gaps between logit matching and sampling from black-box APIs.16:16–19:33 · Guest teaching 6/10 Compute and Cost-Matched Data Generation Dynamics Rishabh shares empirical findings showing that generating filtered data from a smaller model like Gemma 9B beats self-generation from Gemma 27B under compute-matched conditions. Swyx succinctly captures the core intuition that filtering injects external signal.19:36–27:04 · Guest teaching 7/10 Mitigating Exposure Bias via On-Policy Distillation Rishabh uses the DAgger driving analogy to explain exposure bias and train-test distribution shift when students make compounding errors. Swyx immediately identifies the phenomenon as exposure bias as Rishabh outlines the math of on-policy reverse KL distillation.27:06–31:35 · Guest teaching 6/10 Theoretical Divergence: Mode-Covering Versus Mode-Seeking Rishabh presents the theoretical differences between mode-covering forward KL and mode-seeking reverse KL using Gaussian curves. Swyx pushes on the realism of toy Gaussian assumptions in natural language, prompting Rishabh to show real diversity-performance trade-off curves.31:38–34:04 · Guest teaching 5/10 Repurposing RLHF Infrastructure for Model Distillation Rishabh explains how to implement on-policy distillation inside existing RLHF frameworks by zeroing the reward term and retargeting the KL anchor. Swyx directly challenges the framing, pressing Rishabh twice on why turning off rewards is considered profound.34:05–37:44 · Guest teaching 6/10 Speculative Decoding Acceleration and Active Teacher Intervention Rishabh details how distillation accelerates speculative decoding and previews recent research on active teacher intervention during student rollout generation. Swyx synthesizes the takeaway that every large model deployment benefits from a distilled student companion.37:44–44:33 · Guest teaching 7/10 Online Versus Offline Distillation for Complex Reasoning Swyx challenges the practical viability of online distillation, arguing that keeping a live teacher model during training defeats the cost-saving purpose. Rishabh counters firmly with the amortized serving economics of distilling once to serve billions of requests.5:48–8:51 · Guest disagreement 1/10 Dark Knowledge and Post-Training Pipeline Integration Alessio and Swyx explore the high-level motivation for distillation and ask whether post-training pipelines like RLHF can be bypassed. Rishabh explains how distillation fits into the post-training stack and defines Geoff Hinton's concept of dark knowledge.8:54–12:15 · Guest disagreement 0/10 Logit-Based Distillation and Soft Probability Labels Rishabh details classical logit-matching distillation and how soft probability labels supply richer supervision than hard one-hot tokens. Swyx raises a pertinent question regarding compute-matched versus data-matched training FLOPs.12:15–16:13 · Guest disagreement 1/10 Synthetic Data Distillation and Output Verification Rishabh outlines synthetic data distillation and how DeepSeek-R1 utilized best-of-n verifiable sampling across disparate tokenizers. Swyx asks about efficiency gaps between logit matching and sampling from black-box APIs.16:16–19:33 · Guest disagreement 0/10 Compute and Cost-Matched Data Generation Dynamics Rishabh shares empirical findings showing that generating filtered data from a smaller model like Gemma 9B beats self-generation from Gemma 27B under compute-matched conditions. Swyx succinctly captures the core intuition that filtering injects external signal.19:36–27:04 · Guest disagreement 0/10 Mitigating Exposure Bias via On-Policy Distillation Rishabh uses the DAgger driving analogy to explain exposure bias and train-test distribution shift when students make compounding errors. Swyx immediately identifies the phenomenon as exposure bias as Rishabh outlines the math of on-policy reverse KL distillation.27:06–31:35 · Guest disagreement 1/10 Theoretical Divergence: Mode-Covering Versus Mode-Seeking Rishabh presents the theoretical differences between mode-covering forward KL and mode-seeking reverse KL using Gaussian curves. Swyx pushes on the realism of toy Gaussian assumptions in natural language, prompting Rishabh to show real diversity-performance trade-off curves.31:38–34:04 · Guest disagreement 2/10 Repurposing RLHF Infrastructure for Model Distillation Rishabh explains how to implement on-policy distillation inside existing RLHF frameworks by zeroing the reward term and retargeting the KL anchor. Swyx directly challenges the framing, pressing Rishabh twice on why turning off rewards is considered profound.34:05–37:44 · Guest disagreement 0/10 Speculative Decoding Acceleration and Active Teacher Intervention Rishabh details how distillation accelerates speculative decoding and previews recent research on active teacher intervention during student rollout generation. Swyx synthesizes the takeaway that every large model deployment benefits from a distilled student companion.37:44–44:33 · Guest disagreement 4/10 Online Versus Offline Distillation for Complex Reasoning Swyx challenges the practical viability of online distillation, arguing that keeping a live teacher model during training defeats the cost-saving purpose. Rishabh counters firmly with the amortized serving economics of distilling once to serve billions of requests.5:48–8:51 · The hosts pushing back 1/10 Dark Knowledge and Post-Training Pipeline Integration Alessio and Swyx explore the high-level motivation for distillation and ask whether post-training pipelines like RLHF can be bypassed. Rishabh explains how distillation fits into the post-training stack and defines Geoff Hinton's concept of dark knowledge.8:54–12:15 · The hosts pushing back 1/10 Logit-Based Distillation and Soft Probability Labels Rishabh details classical logit-matching distillation and how soft probability labels supply richer supervision than hard one-hot tokens. Swyx raises a pertinent question regarding compute-matched versus data-matched training FLOPs.12:15–16:13 · The hosts pushing back 1/10 Synthetic Data Distillation and Output Verification Rishabh outlines synthetic data distillation and how DeepSeek-R1 utilized best-of-n verifiable sampling across disparate tokenizers. Swyx asks about efficiency gaps between logit matching and sampling from black-box APIs.16:16–19:33 · The hosts pushing back 1/10 Compute and Cost-Matched Data Generation Dynamics Rishabh shares empirical findings showing that generating filtered data from a smaller model like Gemma 9B beats self-generation from Gemma 27B under compute-matched conditions. Swyx succinctly captures the core intuition that filtering injects external signal.19:36–27:04 · The hosts pushing back 1/10 Mitigating Exposure Bias via On-Policy Distillation Rishabh uses the DAgger driving analogy to explain exposure bias and train-test distribution shift when students make compounding errors. Swyx immediately identifies the phenomenon as exposure bias as Rishabh outlines the math of on-policy reverse KL distillation.27:06–31:35 · The hosts pushing back 2/10 Theoretical Divergence: Mode-Covering Versus Mode-Seeking Rishabh presents the theoretical differences between mode-covering forward KL and mode-seeking reverse KL using Gaussian curves. Swyx pushes on the realism of toy Gaussian assumptions in natural language, prompting Rishabh to show real diversity-performance trade-off curves.31:38–34:04 · The hosts pushing back 5/10 Repurposing RLHF Infrastructure for Model Distillation Rishabh explains how to implement on-policy distillation inside existing RLHF frameworks by zeroing the reward term and retargeting the KL anchor. Swyx directly challenges the framing, pressing Rishabh twice on why turning off rewards is considered profound.34:05–37:44 · The hosts pushing back 1/10 Speculative Decoding Acceleration and Active Teacher Intervention Rishabh details how distillation accelerates speculative decoding and previews recent research on active teacher intervention during student rollout generation. Swyx synthesizes the takeaway that every large model deployment benefits from a distilled student companion.37:44–44:33 · The hosts pushing back 5/10 Online Versus Offline Distillation for Complex Reasoning Swyx challenges the practical viability of online distillation, arguing that keeping a live teacher model during training defeats the cost-saving purpose. Rishabh counters firmly with the amortized serving economics of distilling once to serve billions of requests.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 41:09 Rishabh rejects the claim that live teachers defeat the purpose

When Swyx asserts that running a live teacher during training negates distillation advantages, Rishabh immediately rejects the premise with the amortized principle that models are distilled once but served billions of times.

Hardest push from the hosts ▶ 32:30 Swyx questions the profundity of modifying RLHF loops

Swyx interrupts Rishabh to demand clear justification, asking twice why simply turning off reward maximization and adjusting KL anchors in RLHF infra is considered profound.

Biggest teaching moment ▶ 20:25 The DAgger car crash analogy for autoregressive error accumulation

Rishabh clearly educates the hosts on why behavior cloning fails in deployment using the RL driving example where expert teachers never demonstrate recovery from catastrophic drift.

The host holds their own ▶ 5:04 Swyx defines the Pareto frontier slope as the state of distillation

Swyx offers an original conceptual framework equating the slope of the capability-cost Pareto curve to the state of distillation art, which Rishabh praises as an insight he had not considered.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Dark Knowledge and Post-Training Pipeline Integration 5511 Alessio and Swyx explore the high-level motivation for distillation and ask whether post-training pipelines like RLHF can be bypassed. Rishabh explains how distillation fits into the post-training stack and defines Geoff Hinton's concept of dark knowledge.
Logit-Based Distillation and Soft Probability Labels 4601 Rishabh details classical logit-matching distillation and how soft probability labels supply richer supervision than hard one-hot tokens. Swyx raises a pertinent question regarding compute-matched versus data-matched training FLOPs.
Synthetic Data Distillation and Output Verification 4511 Rishabh outlines synthetic data distillation and how DeepSeek-R1 utilized best-of-n verifiable sampling across disparate tokenizers. Swyx asks about efficiency gaps between logit matching and sampling from black-box APIs.
Compute and Cost-Matched Data Generation Dynamics 5601 Rishabh shares empirical findings showing that generating filtered data from a smaller model like Gemma 9B beats self-generation from Gemma 27B under compute-matched conditions. Swyx succinctly captures the core intuition that filtering injects external signal.
Mitigating Exposure Bias via On-Policy Distillation 5701 Rishabh uses the DAgger driving analogy to explain exposure bias and train-test distribution shift when students make compounding errors. Swyx immediately identifies the phenomenon as exposure bias as Rishabh outlines the math of on-policy reverse KL distillation.
Theoretical Divergence: Mode-Covering Versus Mode-Seeking 5612 Rishabh presents the theoretical differences between mode-covering forward KL and mode-seeking reverse KL using Gaussian curves. Swyx pushes on the realism of toy Gaussian assumptions in natural language, prompting Rishabh to show real diversity-performance trade-off curves.
Repurposing RLHF Infrastructure for Model Distillation 6525 Rishabh explains how to implement on-policy distillation inside existing RLHF frameworks by zeroing the reward term and retargeting the KL anchor. Swyx directly challenges the framing, pressing Rishabh twice on why turning off rewards is considered profound.
Speculative Decoding Acceleration and Active Teacher Intervention 4601 Rishabh details how distillation accelerates speculative decoding and previews recent research on active teacher intervention during student rollout generation. Swyx synthesizes the takeaway that every large model deployment benefits from a distilled student companion.
Online Versus Offline Distillation for Complex Reasoning 6745 Swyx challenges the practical viability of online distillation, arguing that keeping a live teacher model during training defeats the cost-saving purpose. Rishabh counters firmly with the amortized serving economics of distilling once to serve billions of requests.

Statements from this episode (24)

Opinion
Swyx: Frontier Models Exist Primarily to Distill Smaller, Usable Models
“Even GPT 4.5 is too expensive. Normally it's really gonna use it in, in any reasonable quantity. Like, you know, Claude 3.5 Opus, like if it does exist, still not like, you know, the thing that we actually use is Sonnet, right? So like, it's almost like a depl…”
Shawn Wang Mar 23, 2025 ▶ 3:58
Insight
Agarwal: Distillation Drives Year-Over-Year AI Capability Cost Reductions
“The capability which we have right now, maybe next year will be much cheaper to have that same thing. And that likely is the result of distillation, right? Like that's, it's not just because we are doing or figured out something magical. It is because distilla…”
Rishabh Agarwal Mar 23, 2025 ▶ 4:47
Insight
Swyx: The Cost-Performance Pareto Frontier Slope Reflects Distillation SOTA
“The slope of the Pareto Frontier is basically the state of the art of distillation, and the gentler the slope, the better distillation is.”
Shawn Wang Mar 23, 2025 ▶ 5:11
Insight
Agarwal: Distilling a Large Model Outperforms Direct Training on the Same Data
“This is something that people have found again and again, that basically you can train a model on some data, or you can train a bigger model on that data and distill that model to another model, and that distill model is better.”
Rishabh Agarwal Mar 23, 2025 ▶ 5:54
Insight
Agarwal: Optimal Post-Training Pipeline Combines Heavy Distillation Followed by RL
“So, so I would think maybe an optimal pipeline would look like you do distillation heavily, but then you still do some RL afterwards, because maybe there's still something you can get out of your reward functions or whatever your post-training stack is.”
Rishabh Agarwal Mar 23, 2025 ▶ 8:05
Assertion Supported
Agarwal: Gemma 2 Used Soft-Label Logit Distillation During Pre-Training
“GemRTool used distillation for pre-training, where they used logits, or these soft labels, which is rather than having hard zero, one tokens, which is, I want to predict this next token, they have like soft labels for all possible tokens.”
Rishabh Agarwal Mar 23, 2025 ▶ 10:05
Assertion Supported
Agarwal: DeepSeek Distilled Reasoning Models Using Correctness-Filtered Synthetic Data
“What they did was they took the best model they had, they generated a bunch of samples, they filtered them based on correctness, because these were on tasks like coding, mathematical problem solving, and a bunch of those things. So they saw whatever, what are …”
Rishabh Agarwal Mar 23, 2025 ▶ 13:57
Insight
Agarwal: Synthetic data distillation bypasses model vocabulary and tokenizer mismatches
“The one nice thing about this kind of distillation is it doesn't matter if you have a vocabulary mismatch, because we're not using the next token distribution or probability labels. You can distill from one model which uses some random tokenizer to another mod…”
Rishabh Agarwal Mar 23, 2025 ▶ 14:26
Opinion
Agarwal: Developers should use logit distillation over synthetic data for open models
“But for open models, yeah, you can do better because you have access to logits, so why not use them? At least that's my take.”
Rishabh Agarwal Mar 23, 2025 ▶ 15:29
Assertion Supported
Agarwal: Filtered 9B Synthetic Data Outperforms 27B Self-Generated Data
“One thing we found consistently, so here what we had two models, nine Gemma, nine B and Gemma, 27 B, and we found consistently that actually generating data from nine B in a compute match setting is always better, even better for distilling or actually improvi…”
Rishabh Agarwal Mar 23, 2025 ▶ 17:41
Insight
Agarwal: Inverting KL Divergence Direction Derives On-Policy Distillation
“All you need to do, at least in math, if you just flip the direction of your KL divergence, so earlier we were minimizing KL between a teacher and student model, if we just swap that order, because KL is not symmetric, you will actually get this kind of distil…”
Rishabh Agarwal Mar 23, 2025 ▶ 24:37
Assertion Supported
Agarwal: Synthetic Data Distillation Can Outperform Logits on Benchmarks
“It's not always the case that synthetic data distillation beats, oh, sorry, it's worse than logits. Like, sometimes logits can be worse off. So if you look at the T-five base, two-fifty million scenario on GSM-HK, On the last plot, you can see that synthetic d…”
Rishabh Agarwal Mar 23, 2025 ▶ 26:20
Disclosure
Agarwal: Google DeepMind Used On-Policy Distillation in Gemma Post-Training
“So this kind of thing was used in Jemma DuPo's training. It's mentioned briefly, like, there's no details there, but it was used there.”
Rishabh Agarwal Mar 23, 2025 ▶ 26:58
Insight
Agarwal: Standard distillation KL loss places mass where teacher has none
“The typically what we use is this mode covering KL, the one on the left. That is the standard distillation loss that everyone uses. But you can see the behavior. And you can already see what's weird about it. It's putting a lot of mass on places where there's …”
Rishabh Agarwal Mar 23, 2025 ▶ 28:53
Insight
Agarwal: Distillation KL direction dictates trade-off between diversity and performance
“There's a trade-off between diversity and performance. So on the y-axis, I'm showing performance. On the x-axis, I'm showing similarity or basically how, like one minus diversity. So more similar things are less diverse. And you can see, depending on the diver…”
Rishabh Agarwal Mar 23, 2025 ▶ 30:24
Insight
Agarwal: 50/50 mix of KL divergences usually works when goals are unclear
“The general recommendation I would give people is that maybe use a mixture of half and half. Like that's what some people have used, right? That's like saying, yeah, basically saying, I don't know what I want. I just want something to work well enough. I'll ju…”
Rishabh Agarwal Mar 23, 2025 ▶ 31:08
Insight
Agarwal: Standard RLHF Infrastructure Can Be Repurposed for Model Distillation
“First step is you go to an RLHF, RLXF, whatever framework you have. You turn off the reward term. So you delete the reward part of it. All you are left is some KL term. And now you swap your KL, the anchor policy, to a bigger teacher policy. And there you go. …”
Rishabh Agarwal Mar 23, 2025 ▶ 31:54
Insight
Agarwal: RLHF and LLM distillation can be combined simultaneously
“You can actually combine RLHF and distillation together, because now you're doing two things at the same time.”
Rishabh Agarwal Mar 23, 2025 ▶ 33:17
Insight
Agarwal: Distillation accelerates speculative decoding for large models
“Now, the thing is, the effectiveness of this method depends on how close the sampler, the small model is to the bigger model that we want to speed up, and actually distillation exactly fixes that, which is, by distillation, you can make things closer to each o…”
Rishabh Agarwal Mar 23, 2025 ▶ 34:57
Assertion Supported
Agarwal: Google AI Overviews Uses Speculative Decoding With Distilled Models
“And this was actually used, so I guess the cool application of this I can mention is the next slide, which is AI overviews at Google. I don't know if people have seen this or this kind of thing. Like, there is this thing that comes up, and that actually uses t…”
Rishabh Agarwal Mar 23, 2025 ▶ 35:46
Assertion Supported
Agarwal: RL literature shows on-policy distillation beats offline methods for agents
“The other thing is in the RL literature, there are, like, results which show that actually this kind of distillation is much more optimal for agentic tasks or really long horizon tasks.”
Rishabh Agarwal Mar 23, 2025 ▶ 38:45
Insight
Agarwal: Synthetic data distillation achieves 80% to 90% of target gains
“Try the simplest thing first, which is synthetic data distillation. That already gets you to 80% of the job or 90% of the job.”
Rishabh Agarwal Mar 23, 2025 ▶ 41:42
Prediction Not checkable as stated
Agarwal: Logit Distillation Can Match Giant Teacher Models on Reasoning
“My hunch is that the logic-based distillation can go even further, and you might be able to even close the gap with the biggest of the teachers you have, because I don't think you need a huge number of parameters, because the reasoning process is very, very, l…”
Rishabh Agarwal Mar 23, 2025 ▶ 42:39
Insight
Agarwal: Classical RL ideas from games are relevant for LLM agents
“When we are again talking about things like agents and reasoning and whatnot, I think a lot of the things are probably done on a small scale in the RL literature, in games and whatnot, and a lot of those ideas would be relevant again now.”
Rishabh Agarwal Mar 23, 2025 ▶ 46:04
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.