REINFORCE
topic on 2 shows · 1 statements across 1 episodes · said 1 times in 1 episodes since 2025
Mentions by year, every show
tap a year for its mentions
No Priors 1
2025 1 mention in 1 episode
every mention on every show, scene by scene, with the transcript →
1 statements about REINFORCE, every show
Jin: Policy gradient algorithms function as weighted supervised fine-tuning
“If you kind of, like, look at, if you kind of stare at, like, this part it sort of looks like just, like, weighted supervised fine-tuning, right? Like, you have this, like, log of, like, the probability of a token and, like, some, like, weight on it and reinfo…”