Reinforcement Learning Policy
topic on 1 show · 1 statements across 1 episodes
1 statements about Reinforcement Learning Policy, every show
Jin: Language models map directly to reinforcement learning policies
“So the states are, like, the text prefixes, so the initial states, like, the prompt the actions are the next tokens that means, like, a language model is, like, exactly what a policy is, right? A policy maps a state to a probability distribution of our next ac…”