Pairwise Preference
topic on 1 show · 1 statements across 1 episodes
1 statements about Pairwise Preference, every show
Lambert: Training LLM reward models on 0-to-10 ratings failed
“People tried that with language models, which is if you have a prompt and a completion and you just have someone rate it from zero to 10, could you then train a reward model on all of these completions and zero to 10 ratings and see if you could actually chang…”