The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 7 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Assertion Supported
Agarwal: Filtered 9B Synthetic Data Outperforms 27B Self-Generated Data
“One thing we found consistently, so here what we had two models, nine Gemma, nine B and Gemma, 27 B, and we found consistently that actually generating data from nine B in a compute match setting is always better, even better for distilling or actually improvi…”
Rishabh Agarwal Mar 23, 2025 ▶ 17:41 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Prediction Not checkable as stated
Agarwal: Logit Distillation Can Match Giant Teacher Models on Reasoning
“My hunch is that the logic-based distillation can go even further, and you might be able to even close the gap with the biggest of the teachers you have, because I don't think you need a huge number of parameters, because the reasoning process is very, very, l…”
Rishabh Agarwal Mar 23, 2025 ▶ 42:39 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Assertion Supported
Agarwal: Gemma 2 Used Soft-Label Logit Distillation During Pre-Training
“GemRTool used distillation for pre-training, where they used logits, or these soft labels, which is rather than having hard zero, one tokens, which is, I want to predict this next token, they have like soft labels for all possible tokens.”
Rishabh Agarwal Mar 23, 2025 ▶ 10:05 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Assertion Supported
Agarwal: DeepSeek Distilled Reasoning Models Using Correctness-Filtered Synthetic Data
“What they did was they took the best model they had, they generated a bunch of samples, they filtered them based on correctness, because these were on tasks like coding, mathematical problem solving, and a bunch of those things. So they saw whatever, what are …”
Rishabh Agarwal Mar 23, 2025 ▶ 13:57 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Assertion Supported
Agarwal: Synthetic Data Distillation Can Outperform Logits on Benchmarks
“It's not always the case that synthetic data distillation beats, oh, sorry, it's worse than logits. Like, sometimes logits can be worse off. So if you look at the T-five base, two-fifty million scenario on GSM-HK, On the last plot, you can see that synthetic d…”
Rishabh Agarwal Mar 23, 2025 ▶ 26:20 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Assertion Supported
Agarwal: Google AI Overviews Uses Speculative Decoding With Distilled Models
“And this was actually used, so I guess the cool application of this I can mention is the next slide, which is AI overviews at Google. I don't know if people have seen this or this kind of thing. Like, there is this thing that comes up, and that actually uses t…”
Rishabh Agarwal Mar 23, 2025 ▶ 35:46 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Assertion Supported
Agarwal: RL literature shows on-policy distillation beats offline methods for agents
“The other thing is in the RL literature, there are, like, results which show that actually this kind of distillation is much more optimal for agentic tasks or really long horizon tasks.”
Rishabh Agarwal Mar 23, 2025 ▶ 38:45 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.