RL Scaling
topic on 2 shows · 3 statements across 3 episodes
the MAD Podcast
Big Technology
3 statements about RL Scaling, every show
Bourgeau: Vibe coding performance stems mostly from RL scaling and post-training
“I think this is, yeah, this is in general for vibe coding specifically, I think that's maybe more of an RL scaling and post training thing where, where you can actually get quite a lot of data and train them all to do that really well.”
Douglas: OpenAI's o1 established test-time compute and RL as a scaling axis
“And I think OpenAI deserves a lot of credit for you know, releasing the first, like, serious RL plus LLMs release with O-one. And I think this really kicked off a pretty, you know, substantial change because it opened up a new axis of scaling, right? There was…”