scaling RL
1 statements across 1 episodes · 1 bullish · 0 bearish · 1 people on the record · first statement Nov 14, 2024 by Rishi Mehta · across every show →
Everything said about scaling RL, oldest first
Nov 14, 2024 bullish
Mehta: AlphaProof's RL scaling and test-time compute generalize across domains
“Some of the sort of tech we developed here of like, you know, like scaling RL and like figuring out how to spend a lot of inference time compute stuff like this feels like it's Quite generally applicable to many other problems.”