RLHF preference datasets
1 statements across 1 episodes · 0 bullish · 0 bearish · 1 people on the record · first statement Jan 11, 2024 by Nathan Lambert · across every show →
Everything said about RLHF preference datasets, oldest first
Jan 11, 2024 neutral
Lambert: RLHF reward models achieve only 65% to 75% validation agreement
“If you look at a test set, you'll have a chosen and rejected, and you can take the reward model you're training, pass in those completions, And you see if the chosen predicted reward, so the scalar number is higher than the rejected predicted reward, and this …”