pairwise preference
1 statements across 1 episodes · 0 bullish · 0 bearish · 1 people on the record · first statement Jan 11, 2024 by Nathan Lambert · across every show →
Everything said about pairwise preference, oldest first
Jan 11, 2024
Lambert: Training LLM reward models on 0-to-10 ratings failed
“People tried that with language models, which is if you have a prompt and a completion and you just have someone rate it from zero to 10, could you then train a reward model on all of these completions and zero to 10 ratings and see if you could actually chang…”