Model Distillation

topic on 8 shows · 14 statements across 14 episodes

More or Less the Neon Show No Priors Invest Like the Best the a16z Podcast Big Technology TBPN 20VC

14 statements about Model Distillation, every show

20VC Assertion Supported
Atallah: Most Chinese open-weight models permit distillation for reinforcement learning
“The nice thing about the open weight models and the Chinese models that they allow distillation and they like most of them. And that means that you can like take the outputs of these models to do reinforcement learning on top of the model that you're building.”
Alex Atallah Aug 9, 2026 ▶ 55:08 OpenRouter CEO: Why Chinese Open Models Are Beating the US | Why Enterprises Fear OpenAI & Anthropic
a16z Opinion
Moe: Model distillation is not the primary driver of Chinese AI progress
“So I really don't think from currently what we're seeing this is a big cornerstone of what's powering the progress today. In the end, what's powering the progress is still just really smart people with very interesting algorithms, data environment, and they wi…”
Simon Moe Aug 5, 2026 ▶ 44:13 How Open Source Became AI's Backbone | Inferact with a16z
TBPN Insight
Coogan: Broad model access is fundamentally at odds with preventing distillation
“The whole idea of, like, if the, if an amazing intelligence, an amazing model is made, should it be widely available? That's actually at odds with distillation, because you have to hunt those bad actors down, and you play this whack-a-mole.”
John Coogan Jul 30, 2026 ▶ 20:09 U.S. Bans New Chinese Humanoids, Zuck’s Op-Ed, eBay's $56M Lawsuit | Diet TBPN
Altman: Competitors Distilling OpenAI Models Is Not in His Top 10 Worries
“I would rather people not to steal from us for sure. Maybe I'm feeling too confident right now about our progress and what's like the models that are coming. But this is not in like my top 10 list of worries.”
Sam Altman Jul 28, 2026 ▶ 14:14 Sam Altman on AGI, Compute, and Human Agency · Invest Like The Best
20VC Assertion Partly supported
O'Driscoll: Chinese AI models rely on distillation illegal in the US
“Maybe it's because the dirty little secret is a lot of their advantage is distillation, which you can't legally do if you're US-based.”
Rory O'Driscoll Jul 23, 2026 ▶ 17:27 Should the US Ban Chinese Open-Source Models | OpenRouter's Chance To Sell | Stripe Buying PayPal
Anthropic Gated Fable to Block Model Distillation Rather Than Misuse
“My hot take is that one of the things that we saw from Anthropic with all these safeguards that it put on Fable was not as much that they were worried that Fable would be misused, although I'm sure there was some worry there. It was more that, like, we know ou…”
Alex Kantrowitz Jun 29, 2026 ▶ 19:26 Mythos is Back, OpenAI Releases GPT 5.6, Apple’s Price Increases
TBPN Prediction Not checkable as stated
Hays: Consumer AI Access Enables Syndicates to Distill Frontier Models
“If you make these models available through everyday consumer subscriptions, there will be groups that will create networks of, you know, thousands of accounts and distill the models. And then you're giving away this, like, advanced capability.”
Jordi Hays Jun 27, 2026 ▶ 28:37 Shifts In The Creator Economy, Kylie Jenner x Meta, GPT 5.6 Limited Release | Diet TBPN
NO PRIORS Opinion
Helberg: Regulating model distillation is critical to protecting AI investments
“The whole model distillation debate, which is super important to actually protect the economic value of these hundreds of billions of dollars in investments in AI companies.”
Jacob Helberg May 14, 2026 ▶ 30:27 Pax Silica: Inside the Trump Administration’s Tech Strategy with Jacob Helberg
NEON SHOW Insight
Kamath: Distilled AI models match large model accuracy at a fraction of size
“When you use that pre-training model as a teacher for this small model, and we could go into what teacher means, but it would basically Act at similar accuracies as the larger model at 100 or one 10th of the size.”
Sudarshan Kamath Mar 6, 2026 ▶ 6:11 Where SMALL models will Win | Sudarshan kamath, Smallest ai
a16z Insight
Midha: Distilling frontier capabilities from US AI outputs is easy
“No, actually it's not that hard to distill on the outputs of our labs.”
Anjney Midha Aug 15, 2025 ▶ 9:58 The Current Reality of American AI Policy: From ‘Pause AI’ to ‘Build’
20VC Assertion Not checkable as stated
Krieger: AI Labs Use Internal Distillation to Reduce Latency and Costs
“Even, like, let's take within the labs, like, I assume every single one of the labs is using, like, even within themselves, like, it is very valuable to be able to take, you know, the knowledge of your highest-end model and then be able to make it higher, you …”
Mike Krieger Mar 3, 2025 ▶ 28:44 Mike Krieger, Instagram CoFounder & Anthropic CPO: Where Will Value Be Created in an AI World?|E1265 · 20VC with Harry Stebbings
Sam Lessin: AI labs hypocritically protest distillation after scraping the web
“That like you have this thing where like, oh, you want to have it both ways. So you think it's okay to break the economics of the underlying internet you're leveraging, but then you're mad that someone's doing it to you.”
Sam Lessin Jan 31, 2025 ▶ 16:48 #84: China's DeepSeek Shakes the Tech World · More or Less Podcast
NO PRIORS Prediction Not checkable as stated
Karpathy: A 1-billion parameter model will suffice as a cognitive core
“I think even a 1,000,000,001, billion suffices. We'll probably get to that point, and the models can be very, very small. And I think the reason they can be very small is fundamentally, I think, just, like, distillation works.”
Andrej Karpathy Sep 5, 2024 ▶ 27:46 No Priors Ep. 80 | With Andrej Karpathy from OpenAI and Tesla
a16z Insight
Dolgov: Distilling huge AI models beats training small models directly
“Where you're much better off training a huge model and then distilling it into a smaller model than just training small models.”
Dmitry Dolgov Aug 5, 2024 ▶ 14:37 How Waymo Is Using GenAI to Build a Better Driver

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.