RLHF
5 statements across 5 episodes · 3 bullish · 0 bearish · 5 people on the record · first statement Nov 8, 2023 by Sharon Zhou · across every show →
Everything said about RLHF, oldest first
Nov 8, 2023 positive
Zhou: Zero-shot LLMs can replace manual human labeling in RLHF workflows
“Which is that why can't it be another LLM or a pipeline of LLMs that can help with that feedback? I think manual labeling is very tedious, especially for our target user, which is a software engineer. And I don't think people should necessarily have to do all …”
Dec 20, 2023 positive
Nov 20, 2025 positive
Lambert: RLVR targets performance characteristics better than traditional RLHF reward models
“These reward models tend to have a lot of problems and you can over optimize them much more easily because the reward models will pick up on features that are maybe emojis or something like this that you don't actually care about where RLVR is much better matc…”
Nov 26, 2025 neutral
Aug 6, 2026 neutral
Wolf: Frontier AI Training Has Shifted From RLHF to Pure RL
“What we know though, is we moved from this pure, like human data, you know, that was first just pre-training on human data and then also aligning with like human preferences that was called RLHF, where we had a lot of human in the loop and human data. To like …”