reinforcement learning policy
1 statements across 1 episodes · 0 bullish · 0 bearish · 1 people on the record · first statement Apr 29, 2025 by Roger Jin · across every show →
Everything said about reinforcement learning policy, oldest first
Apr 29, 2025
Jin: Language models map directly to reinforcement learning policies
“So the states are, like, the text prefixes, so the initial states, like, the prompt the actions are the next tokens that means, like, a language model is, like, exactly what a policy is, right? A policy maps a state to a probability distribution of our next ac…”