RLHF
37 statements across 10 episodes · 14 bullish · 8 bearish · 9 people on the record · first statement Jun 20, 2023 by George Hotz · across every show →
Everything said about RLHF, oldest first
Jun 20, 2023 negative
Jan 11, 2024 neutral
Jan 11, 2024 negative
Lambert: Startups should avoid RLHF unless it offers niche advantage
“I don't really recommend most startups to do it unless it's like going to provide them a clear competitive advantage in their kind of niche. Yeah. Because you're not going to make your model ChatGPT like better than OpenAI or anything like that.”
Jan 11, 2024 neutral
Lambert: AI field hasn't learned much more, just articulates ignorance better
“I think it's, I try to do it every six or 12 months is my current, is my estimated cadence, just to refine the ways that I say things, and people will see That we don't know that much more, but we have a bit of better way of saying what we don't know.”
Jan 11, 2024 negative
Jan 11, 2024 neutral
Lambert: RLHF avoids controversial preferences, focusing on correctness and style
“The reason this really is done on a deep level is that you're not actually trying to model any, like, contestable preference in this. Like, you're not trying to go into things that are controversial or anything. It's really, like, the notion of preference is t…”
Jan 11, 2024 positive
Lambert: Instruction tuning is more important than RLHF for most practitioners
“I think for most people, instruction tuning is probably still more important in their day-to-day life. I think instruction tuning works very well. You can write samples by hand that make sense. You can get the model to learn from them. You could do this with v…”
Jan 11, 2024 positive
Jan 11, 2024 neutral
Jan 11, 2024 negative
Jan 11, 2024 negative
Lambert: Reinforcement Learning Does Far Less for Alignment Than People Think
“It's like, now you're getting to the point where you don't even really need this to get a good model, so that's why it's like, okay, the RL is such a small part of the actual, like, doing RLHF. Like, RLHF is a metaphor for, like, all language model adaptation,…”
Jan 11, 2024
Jan 11, 2024 neutral
Lambert: Real-Time Human Feedback in LLM RL Loops Is Far From Feasible
“Setting up the infrastructure to take tens of thousands of prompts and generate them and then show them to a human and collect the human responses and then show that Shove that into your training architecture is very far away from working, so we don't really h…”
Jan 11, 2024 neutral
Lambert: RLHF operates as a contextual bandit problem, not sequential RL
“What happens is the state is a prompt, and then you do a completion, and then you throw it away, and you grab a new prompt. Where really in like RL, you, as an RL researcher, you would think of this as being like, you take a state, you complete, Get some compl…”
Jan 11, 2024 bullish
Jan 11, 2024 neutral
Jan 11, 2024 negative
Lambert: Frontier AI labs do not recruit economics or social choice academics
“The RLHF techniques that people use were built in, like, labs like OpenAI and DeepMind, where there are some of these people, they have, they, these places do a pretty good job of trying to get these people in the door when you compare them to, like, startups,…”
Jan 11, 2024 neutral
Lambert: RLHF Performance Depends on Data and Systems Over RL Details
“It really ends up being kind of, like, gibberish that I think is less important now, because it's more about data and infrastructure than RL details, than, like, value functions and everything. A lot of the papers have different terms in the equations. I think…”
Jan 11, 2024 bullish
Lambert: Offline RL could take off due to simpler training pipelines
“There's a few papers that people have published
Not a lot of traction.
I think it could take off.
Some people that I know in the RLHF area really think a lot of people are doing this in industry just because it makes the kind of training process simpler and th…”
Jan 11, 2024 positive
Jan 11, 2024 neutral
Lambert: OpenAI retrains reward models with curated and user prompt mixtures
“And this is like a sort of outer loop optimization that no one in the open is even remotely qualified to talk about, but OpenAI does monitor and they'll like rerun RLHF and train a new reward model with a mixture of their curated data and user prompts to try t…”
Jan 11, 2024 neutral
Jan 11, 2024 neutral
Lambert: RLHF has not been shown to improve underlying model benchmark capabilities
“RLHF is not that shown to improve capabilities yet. I think one of the fun ones is from the GPT-IV technical report. They essentially listed their kind of bogus evaluations, because it's a hilarious table, because it's like LSAT AP exams, and then like AMC-X a…”
Jan 11, 2024 positive
Jan 11, 2024 positive
Lambert: RLHF provides a richer signal per unit of compute than pre-training
“As reinforcement learning is so much less compute, like it, Is a richer signal in terms of its impact, because if they could do what RLHF is doing at pre-training, they would, but they don't know how to have that effect in like a stable manner. Otherwise, ever…”
Jan 11, 2024 negative
Jul 23, 2024 positive
Scialom: RLHF yields superhuman models because humans judge better than they generate
“And because of that, you can have a model that flats the bad outputs, and learns to only shift towards the best and better and better outputs. And you can even end to superhuman abilities, since that I'm bad at writing a poem, but I'm good at judging which one…”
Jul 23, 2024
Jul 23, 2024 positive
Scialom: Smaller Llama 3 models improved via distillation from 405B
“Having bigger models enables to collect better data, for instance, at RLHF stage, because that's the model we use for the annotation. And so we distillate straight forward, like those annotations from this better model to the other models. So I can guarantee y…”
Mar 23, 2025 positive
Agarwal: Standard RLHF Infrastructure Can Be Repurposed for Model Distillation
“First step is you go to an RLHF, RLXF, whatever framework you have. You turn off the reward term. So you delete the reward part of it. All you are left is some KL term. And now you swap your KL, the anchor policy, to a bigger teacher policy. And there you go. …”
Mar 23, 2025 positive
May 23, 2025 positive
Brown: GRPO is more memory efficient and easier to distribute
“GRPO is, like, great for, like, leaning heavy on highly parallel inference compute. It's more memory efficient for the actual training process. It's much easier to do in a distributed fashion because you have less gradient syncing and less model weight copies.”
Jul 31, 2025
Dec 18, 2025 bullish
Zhang: Superhuman computer vision requires RLHF rather than human SFT data
“But if you only do SFT and the SFT data is annotated by human, then your performance is funded by human. You cannot get, kind of, superhuman performance just by, kind of, this kind of data engine approach to use human annotated data and then learn from that. Y…”
Dec 31, 2025
Jan 28, 2026 negative
White: RLHF fails on scientific hypotheses by ignoring impact and information gain
“We learned a lot about how bad our LHF is with people, just like people pay really attention to the tone, to the details, to like how many specific facts or figures on the hypothesis, right? Like actionability about like if the experiment is feasible, but what…”
Aug 15, 2026 positive
Krentsel: Harness architecture must enforce agent rules instead of model alignment
“My bet is this though, that people keep trying to get models to do things that align with their goals and alignment is an unsolved problem. We keep trying to like RLHF, like try to align these models to do the right thing. In harness space, we actually have an…”