RLHF Preference Datasets
topic on 1 show · 1 statements across 1 episodes
1 statements about RLHF Preference Datasets, every show
Lambert: RLHF reward models achieve only 65% to 75% validation agreement
“If you look at a test set, you'll have a chosen and rejected, and you can take the reward model you're training, pass in those completions, And you see if the chosen predicted reward, so the scalar number is higher than the rejected predicted reward, and this …”