RLHF

37 statements across 10 episodes · 14 bullish · 8 bearish · 9 people on the record · first statement Jun 20, 2023 by George Hotz · across every show →

Everything said about RLHF, oldest first

Jun 20, 2023 negative
Opinion
Hotz: RLHF models adopt customer support personalities
“I don't like the RLHF models. I don't like the tuned versions of them. I think that they become, you take on the personality of a customer support agent, right?”
George Hotz Jun 20, 2023 ▶ 1:06:11 Ep 18: Petaflops to the People — with George Hotz of tinycorp
Jan 11, 2024 neutral
Insight
Lambert: RLHF fails completely without a KL divergence constraint
“The KL constraint prevents that. There's not that much documented work on that, but there's a lot of people that know if you take that away, it just doesn't work at all.”
Nathan Lambert Jan 11, 2024 ▶ 35:51 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 negative
Insight
Lambert: Startups should avoid RLHF unless it offers niche advantage
“I don't really recommend most startups to do it unless it's like going to provide them a clear competitive advantage in their kind of niche. Yeah. Because you're not going to make your model ChatGPT like better than OpenAI or anything like that.”
Nathan Lambert Jan 11, 2024 ▶ 29:37 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 neutral
Insight
Lambert: AI field hasn't learned much more, just articulates ignorance better
“I think it's, I try to do it every six or 12 months is my current, is my estimated cadence, just to refine the ways that I say things, and people will see That we don't know that much more, but we have a bit of better way of saying what we don't know.”
Nathan Lambert Jan 11, 2024 ▶ 3:58 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 negative
Insight
Lambert: RLHF reward models lack inductive biases for true preferences
“In RLHF calling things a preference model is a little annoying because There's no inductive bias of what a preference is.”
Nathan Lambert Jan 11, 2024 ▶ 12:41 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 neutral
Insight
Lambert: RLHF avoids controversial preferences, focusing on correctness and style
“The reason this really is done on a deep level is that you're not actually trying to model any, like, contestable preference in this. Like, you're not trying to go into things that are controversial or anything. It's really, like, the notion of preference is t…”
Nathan Lambert Jan 11, 2024 ▶ 39:02 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 positive
Insight
Lambert: Instruction tuning is more important than RLHF for most practitioners
“I think for most people, instruction tuning is probably still more important in their day-to-day life. I think instruction tuning works very well. You can write samples by hand that make sense. You can get the model to learn from them. You could do this with v…”
Nathan Lambert Jan 11, 2024 ▶ 28:11 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 positive
Assertion Not checkable as stated
Lambert: Zephyr was the first open model to succeed with public RLHF
“I think Zephyr was the first model that showed success with RLHF in the public, but that's a long time from everyone knowing that it was something that people are interested in to having any, like, check mark.”
Nathan Lambert Jan 11, 2024 ▶ 51:42 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 neutral
Opinion
Lambert: Frontier Labs Lack Visibility into Cross-Model RLHF Sensitivity
“I think big labs are so over-indexed, are indexed on their own base models, so they don't know, like, what's swapping between CloudBase or GPT-IV-Base, how that would change any notion of preference or what you do with RLHF.”
Nathan Lambert Jan 11, 2024 ▶ 1:30:15 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 negative
Opinion
Lambert: Scaling PPO in RLHF is a Nightmare
“PPO is kind of a nightmare to scale.”
Nathan Lambert Jan 11, 2024 ▶ 1:01:41 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 negative
Insight
Lambert: Reinforcement Learning Does Far Less for Alignment Than People Think
“It's like, now you're getting to the point where you don't even really need this to get a good model, so that's why it's like, okay, the RL is such a small part of the actual, like, doing RLHF. Like, RLHF is a metaphor for, like, all language model adaptation,…”
Nathan Lambert Jan 11, 2024 ▶ 58:36 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024
Insight
Lambert: DPO is closer to RLHF than RLHF is to RL
“I think DPO is closer to RLHF than RLHF is to RL.”
Nathan Lambert Jan 11, 2024 ▶ 1:12:56 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 neutral
Assertion Not checkable as stated
Lambert: Real-Time Human Feedback in LLM RL Loops Is Far From Feasible
“Setting up the infrastructure to take tens of thousands of prompts and generate them and then show them to a human and collect the human responses and then show that Shove that into your training architecture is very far away from working, so we don't really h…”
Nathan Lambert Jan 11, 2024 ▶ 16:13 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 neutral
Insight
Lambert: RLHF operates as a contextual bandit problem, not sequential RL
“What happens is the state is a prompt, and then you do a completion, and then you throw it away, and you grab a new prompt. Where really in like RL, you, as an RL researcher, you would think of this as being like, you take a state, you complete, Get some compl…”
Nathan Lambert Jan 11, 2024 ▶ 16:40 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 bullish
Opinion
Lambert: Major tech companies now realize they need dedicated RLHF teams
“I think any major company that wasn't doing RLHF is now realizing they have to have a team around this.”
Nathan Lambert Jan 11, 2024 ▶ 5:31 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 neutral
Assertion Supported
Lambert: Most open-source RLHF training runs only last a few epochs
“Most RLHF is only a few epochs, at least in the open models”
Nathan Lambert Jan 11, 2024 ▶ 1:09:04 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 negative
Assertion Not checkable as stated
Lambert: Frontier AI labs do not recruit economics or social choice academics
“The RLHF techniques that people use were built in, like, labs like OpenAI and DeepMind, where there are some of these people, they have, they, these places do a pretty good job of trying to get these people in the door when you compare them to, like, startups,…”
Nathan Lambert Jan 11, 2024 ▶ 14:06 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 neutral
Insight
Lambert: RLHF Performance Depends on Data and Systems Over RL Details
“It really ends up being kind of, like, gibberish that I think is less important now, because it's more about data and infrastructure than RL details, than, like, value functions and everything. A lot of the papers have different terms in the equations. I think…”
Nathan Lambert Jan 11, 2024 ▶ 57:58 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 bullish
Prediction Not checkable as stated
Lambert: Offline RL could take off due to simpler training pipelines
“There's a few papers that people have published Not a lot of traction. I think it could take off. Some people that I know in the RLHF area really think a lot of people are doing this in industry just because it makes the kind of training process simpler and th…”
Nathan Lambert Jan 11, 2024 ▶ 1:04:08 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 positive
Insight
Lambert: A 100x smaller language model filters output better than RLHF
“You could use like a hundred times smaller language model and do much better at filtering than RLHF”
Nathan Lambert Jan 11, 2024 ▶ 45:20 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 neutral
Assertion Supported
Lambert: OpenAI retrains reward models with curated and user prompt mixtures
“And this is like a sort of outer loop optimization that no one in the open is even remotely qualified to talk about, but OpenAI does monitor and they'll like rerun RLHF and train a new reward model with a mixture of their curated data and user prompts to try t…”
Nathan Lambert Jan 11, 2024 ▶ 1:32:27 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 neutral
Assertion Not checkable as stated
Lambert: Open-source and research communities lack large-scale RLHF capabilities
“At the same time, we don't have a lot of that in the like open and research communities at the same scale.”
Nathan Lambert Jan 11, 2024 ▶ 5:38 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 neutral
Assertion Supported
Lambert: RLHF has not been shown to improve underlying model benchmark capabilities
“RLHF is not that shown to improve capabilities yet. I think one of the fun ones is from the GPT-IV technical report. They essentially listed their kind of bogus evaluations, because it's a hilarious table, because it's like LSAT AP exams, and then like AMC-X a…”
Nathan Lambert Jan 11, 2024 ▶ 59:53 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 positive
Opinion
Lambert: Anthropic is considered the master of RLHF techniques
“Everyone knows Anthropik is kind of the masters of this”
Nathan Lambert Jan 11, 2024 ▶ 5:49 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 positive
Insight
Lambert: RLHF provides a richer signal per unit of compute than pre-training
“As reinforcement learning is so much less compute, like it, Is a richer signal in terms of its impact, because if they could do what RLHF is doing at pre-training, they would, but they don't know how to have that effect in like a stable manner. Otherwise, ever…”
Nathan Lambert Jan 11, 2024 ▶ 22:14 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 negative
Opinion
Lambert: Annotator disagreement data in RLHF is not used for anything
“I don't really think this disagreement data is used for anything, but it's good to know, like, what the distribution of prompts is, who's doing it, how many samples you have, controlling the workforce.”
Nathan Lambert Jan 11, 2024 ▶ 42:20 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jul 23, 2024 positive
Insight
Scialom: RLHF yields superhuman models because humans judge better than they generate
“And because of that, you can have a model that flats the bad outputs, and learns to only shift towards the best and better and better outputs. And you can even end to superhuman abilities, since that I'm bad at writing a poem, but I'm good at judging which one…”
Thomas Scialom Jul 23, 2024 ▶ 31:56 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Jul 23, 2024
Assertion Supported
Scialom: Meta had to reinvent scaling RLHF without published frontier research
“You have just the basics, but then when it comes to, like, ChatGPT or GPT Instruct or Cloud, No one published the details there. And so we had to reinvent the wheel there in a very short amount of time.”
Thomas Scialom Jul 23, 2024 ▶ 7:53 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Jul 23, 2024 positive
Disclosure
Scialom: Smaller Llama 3 models improved via distillation from 405B
“Having bigger models enables to collect better data, for instance, at RLHF stage, because that's the model we use for the annotation. And so we distillate straight forward, like those annotations from this better model to the other models. So I can guarantee y…”
Thomas Scialom Jul 23, 2024 ▶ 14:02 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Mar 23, 2025 positive
Insight
Agarwal: Standard RLHF Infrastructure Can Be Repurposed for Model Distillation
“First step is you go to an RLHF, RLXF, whatever framework you have. You turn off the reward term. So you delete the reward part of it. All you are left is some KL term. And now you swap your KL, the anchor policy, to a bigger teacher policy. And there you go. …”
Rishabh Agarwal Mar 23, 2025 ▶ 31:54 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Mar 23, 2025 positive
Insight
Agarwal: RLHF and LLM distillation can be combined simultaneously
“You can actually combine RLHF and distillation together, because now you're doing two things at the same time.”
Rishabh Agarwal Mar 23, 2025 ▶ 33:17 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
May 23, 2025 positive
Assertion Supported
Brown: GRPO is more memory efficient and easier to distribute
“GRPO is, like, great for, like, leaning heavy on highly parallel inference compute. It's more memory efficient for the actual training process. It's much easier to do in a distributed fashion because you have less gradient syncing and less model weight copies.”
Will Brown May 23, 2025 ▶ 32:07 ⚡️Multi-Turn RL for Multi-Hour Agents — with Will Brown, Prime Intellect
Jul 31, 2025
Insight
Lambert: RLHF can never be permanently solved
“In the same way that chatbot arena can never be saturated. RLHF can never be solved.”
Nathan Lambert Jul 31, 2025 ▶ 16:42 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Dec 18, 2025 bullish
Insight
Zhang: Superhuman computer vision requires RLHF rather than human SFT data
“But if you only do SFT and the SFT data is annotated by human, then your performance is funded by human. You cannot get, kind of, superhuman performance just by, kind of, this kind of data engine approach to use human annotated data and then learn from that. Y…”
Pengchuan Zhang Dec 18, 2025 ▶ 44:33 SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow)
Dec 31, 2025
Insight
McGrath: RLHF and RLVR differ by data quality, not optimization math
“Really, at the end of the day, like, RLHF, RLVR, They're both policy gradient methods, but the, what's different is just like the input data.”
Josh McGrath Dec 31, 2025 ▶ 9:02 [State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI
Jan 28, 2026 negative
Insight
White: RLHF fails on scientific hypotheses by ignoring impact and information gain
“We learned a lot about how bad our LHF is with people, just like people pay really attention to the tone, to the details, to like how many specific facts or figures on the hypothesis, right? Like actionability about like if the experiment is feasible, but what…”
Andrew White Jan 28, 2026 ▶ 20:35 🔬 From Red Teaming GPT-4 to Automating Drug Discovery: The Future of AI in Science — Andrew White
Aug 15, 2026 positive
Insight
Krentsel: Harness architecture must enforce agent rules instead of model alignment
“My bet is this though, that people keep trying to get models to do things that align with their goals and alignment is an unsolved problem. We keep trying to like RLHF, like try to align these models to do the right thing. In harness space, we actually have an…”
Alex Krentsel Aug 15, 2026 ▶ 14:23 Exo: Harnesses should see their own code and logs — Alex Krentsel, UC Berekeley / Google Research
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.