bandits problem
1 statements across 1 episodes · 0 bullish · 0 bearish · 1 people on the record · first statement Jan 11, 2024 by Nathan Lambert · across every show →
Everything said about bandits problem, oldest first
Jan 11, 2024 neutral
Lambert: RLHF operates as a contextual bandit problem, not sequential RL
“What happens is the state is a prompt, and then you do a completion, and then you throw it away, and you grab a new prompt. Where really in like RL, you, as an RL researcher, you would think of this as being like, you take a state, you complete, Get some compl…”