post-training

also referred to as: post training

7 statements across 6 episodes · 3 bullish · 1 bearish · 6 people on the record · first statement Mar 23, 2025 by Rishabh Agarwal · across every show →

Everything said about post-training, oldest first

Mar 23, 2025 positive
Insight
Agarwal: Optimal Post-Training Pipeline Combines Heavy Distillation Followed by RL
“So, so I would think maybe an optimal pipeline would look like you do distillation heavily, but then you still do some RL afterwards, because maybe there's still something you can get out of your reward functions or whatever your post-training stack is.”
Rishabh Agarwal Mar 23, 2025 ▶ 8:05 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Jun 19, 2025 positive
Disclosure
Brown: OpenAI models undergo mid-training and post-training before release
“For open AI models, like, they go through a mid-training step, and then they go through a post-training step, and then they're released, and they're a lot more useful. Like, frankly, if you interacted with the only pre-trained model, it would be super difficul…”
Noam Brown Jun 19, 2025 ▶ 1:11:38 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Aug 29, 2025 positive
Insight
Morcos: Post-training techniques are better applied in pre- and mid-training
“Most of what we do in post-training is better were done in pre and mid training and earlier on in training in general.”
Ari Morcos Aug 29, 2025 ▶ 42:16 Better Data is All You Need — Ari Morcos, Datology
Aug 29, 2025 negative
Insight
Morcos: Post-training alignment is ineffective long-term compared to pre-training alignment
“Like fundamentally, I think alignment and post training doesn't really make sense as a long-term solution. If you can easily align a model through post training, you can easily misalign a model through post training. If it's easy to put it in, it's easy to tak…”
Ari Morcos Aug 29, 2025 ▶ 53:37 Better Data is All You Need — Ari Morcos, Datology
Jan 23, 2026 neutral
Insight
Yi Tay: 'Reasoning' Technically Just Means Post-Training RL with Thinking Trajectories
“So I think the actual, like, technical definition of reasoning is making models better with thinking and post-training. Ok? Yeah. So basically, like, RL-ing the model to think better.”
Yi Tay Jan 23, 2026 ▶ 33:24 Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay
Jul 8, 2026
Disclosure
Bubna: Modal multi-node training targets post-training, not large-scale pre-training
“And we're not going for obviously like large scale pre-training runs. The thing that we've built multi-handle training for is we see a lot of smaller scale post-training like people are post-training like medium-sized fun models so they can get higher quality …”
Akshat Bubna Jul 8, 2026 ▶ 34:14 The Future of AI Infra: from Kubernetes to Agent Sandboxes — Akshat Bubna, Modal CTO
Jul 22, 2026
Insight
Kant: Big model post-training recipes transfer down to small models, not up
“It's not very helpful to have a post training recipe for a smaller model and try to apply it to a bigger model. Yeah. It just, in all cases, you're gonna have to rethink most of the recipe. But recipe for post training for a bigger model applied to a smaller m…”
Eiso Kant Jul 22, 2026 ▶ 1:44:06 The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.