Jin: Fusing inference and scoring natively enables multi-step and multi-agent RL
Roger Jin · What is an RL environment? w/ Nous Research's Roger Jin · Apr 29, 2025 · at 10:46
This episode carries Roger Jin's own address, with nobody on the show putting questions to them. It still counts as said, and it is kept out of every score on their page.
Roger Jin of Nous Research discusses the design trade-offs of reinforcement learning environment abstractions at ICLR.
“Collect trajectories is a fusion of both these. It handles both inference and scoring, and we deliberately chose that because, like, what happens when you try to, like, do, like, multi-turn, or, like, multi-agent with, like, a separate score function? Then things kind of get, like, weird. So, we have this, like, fused inference and scoring function, collect trajectories, where, we're kind of agnostic to how many, like, How many times, like, inference and scoring happen? So multi-step, multi-agent are kind of, like, automatically supported by this method.”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →