REINFORCE
1 statements across 1 episodes · 0 bullish · 0 bearish · 1 people on the record · first statement Apr 29, 2025 by Roger Jin · across every show →
Everything said about REINFORCE, oldest first
Apr 29, 2025
Jin: Policy gradient algorithms function as weighted supervised fine-tuning
“If you kind of, like, look at, if you kind of stare at, like, this part it sort of looks like just, like, weighted supervised fine-tuning, right? Like, you have this, like, log of, like, the probability of a token and, like, some, like, weight on it and reinfo…”