Jin: Policy gradient algorithms function as weighted supervised fine-tuning
Roger Jin · What is an RL environment? w/ Nous Research's Roger Jin · Apr 29, 2025 · at 5:59
This episode carries Roger Jin's own address, with nobody on the show putting questions to them. It still counts as said, and it is kept out of every score on their page.
Roger Jin of Nous Research connects policy gradient reinforcement learning algorithms such as REINFORCE and GRPO directly to supervised fine-tuning.
“If you kind of, like, look at, if you kind of stare at, like, this part it sort of looks like just, like, weighted supervised fine-tuning, right? Like, you have this, like, log of, like, the probability of a token and, like, some, like, weight on it and reinforces, like, the return, but it doesn't have to be that, like, you can replace it with, like, some advantage like, GRPO does that, right?”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →