multi-turn reinforcement learning
1 statements across 1 episodes · 1 bullish · 0 bearish · 1 people on the record · first statement May 21, 2025 by Silas Alberti · across every show →
Everything said about multi-turn reinforcement learning, oldest first
May 21, 2025 positive
Alberti: Multi-Turn RL Enables Aggressive Code Optimization Over Single-Turn Models
“Basically the single-turn model that was just trained on, like, getting the best result after one turn. It would basically be a little bit, like, too careful, because it couldn't risk writing, like, non-compiling code, whereas, like, the multi-turn model would…”