Human Feedback
topic on 3 shows · 6 statements across 4 episodes
6 statements about Human Feedback, every show
Fitzpatrick: Belief that synthetic data replaces human feedback is wrong
“Look, I think the biggest one is just the view that synthetic data will take over, and you just will not need human feedback.”
Fitzpatrick: Human feedback in AI training will remain essential for 10 years
“For a multi-stage reasoning test that requires a PhD in multi-different languages, and, like, human feedback is going to be important in that for the next decade.”
Nair: RLHF Is a Side Branch Because Compute Cannot Be Scaled
“I think human feedback is kind of like a bit of like a side branch, because you can't really pour that much compute Into it, right? It's like, you take the model, and you, like, elicit it to be a little bit better in terms of personality”
Mann: Scaling models makes finding qualified human evaluators increasingly difficult
“As we've trained the models more and scaled up a lot, it's become harder to find humans with enough expertise to meaningfully contribute to these feedback comparisons. So for example, for coding, somebody who isn't already an expert software engineer would pro…”
Agarwal: AI teams rarely use human feedback, despite high satisfaction when used
“I think it's not used as much. I agree. And I think that's something I tell every customer that you need to close the loop and feedback is going to help you. I think the teams that are doing it are really happy. They're not doing it.”
Agarwal: LLM teams will adopt human feedback only after stabilizing applications
“I think the output side of things will come. Once we've reached stable state, where there are production applications that are stable, and then they want to start optimizing, is when they'll start looking at human feedback as well.”