Scaling RL
topic on 1 show · 1 statements across 1 episodes
1 statements about Scaling RL, every show
Mehta: AlphaProof's RL scaling and test-time compute generalize across domains
“Some of the sort of tech we developed here of like, you know, like scaling RL and like figuring out how to spend a lot of inference time compute stuff like this feels like it's Quite generally applicable to many other problems.”