long horizon RL
1 statements across 1 episodes · 0 bullish · 0 bearish · 1 people on the record · first statement May 15, 2024 by Dwarkesh Patel · across every show →
Everything said about long horizon RL, oldest first
May 15, 2024
Long-horizon reinforcement learning for AI agents is constrained by sparse rewards
“People have been talking about long horizon RL, which is the training method. You need to get something like this, where you go tell it to do something and then you reward it at the end for having achieved that outcome. But the difficulty with those kinds of a…”