RLHF

topic on 12 shows · 61 statements across 33 episodes

the Y Combinator Startup Podcast Innovators & Investors Cheeky Pint We Live to Build Latent Space Lenny's Podcast No Priors A Product Market Fit Show Sourcery the MAD Podcast the a16z Podcast 20VC

The latest 60 statements about RLHF, every show

Krentsel: Harness architecture must enforce agent rules instead of model alignment
“My bet is this though, that people keep trying to get models to do things that align with their goals and alignment is an unsolved problem. We keep trying to like RLHF, like try to align these models to do the right thing. In harness space, we actually have an…”
Alex Krentsel Aug 15, 2026 ▶ 14:23 Exo: Harnesses should see their own code and logs — Alex Krentsel, UC Berekeley / Google Research
MAD Assertion Supported
Wolf: Frontier AI Training Has Shifted From RLHF to Pure RL
“What we know though, is we moved from this pure, like human data, you know, that was first just pre-training on human data and then also aligning with like human preferences that was called RLHF, where we had a lot of human in the loop and human data. To like …”
Thomas Wolf Aug 6, 2026 ▶ 29:08 “OpenAI’s Model Hacked Us” - Hugging Face’s Thomas Wolf
Y COMBINATOR Prediction Not checkable as stated
Hassabis: Current AI paradigms will be part of final AGI architecture
“The components that you just mentioned, I'm pretty sure will be part of the final architecture for AGI. So I think they've come such a long way now and we've proven out so many things about what they can do. I can't see a world in which we will sort of realize…”
Demis Hassabis Apr 29, 2026 ▶ 2:22 Demis Hassabis: Agents, AGI & The Next Big Scientific Breakthrough · Y Combinator
CHEEKY PINT Disclosure
Pichai: Google held back LaMDA due to lack of RLHF and toxicity
“In fact, in the Google I.O. In, in 22, We launched something called AI Test Kitchen, and that was Lambda, but we had constrained it because internally we didn't have an end-to-end version which was RLHFed, right? So the version I saw was a lot more you know, t…”
Sundar Pichai Apr 7, 2026 ▶ 2:36 The history and future of AI at Google, with Sundar Pichai
White: RLHF fails on scientific hypotheses by ignoring impact and information gain
“We learned a lot about how bad our LHF is with people, just like people pay really attention to the tone, to the details, to like how many specific facts or figures on the hypothesis, right? Like actionability about like if the experiment is feasible, but what…”
Andrew White Jan 28, 2026 ▶ 20:35 🔬 From Red Teaming GPT-4 to Automating Drug Discovery: The Future of AI in Science — Andrew White
20VC Insight
Fitzpatrick: RLHF is the only way to accurately fine-tune context-specific AI
“And so the only way to actually do the fine-tuning process consistently And to get it accurate for any specific context is RLHF.”
Matt Fitzpatrick Dec 31, 2025 ▶ 52:12 Matt Fitzpatrick: Who Wins the Data Labelling Race & Why Al Needs Forward-Deployed Engineers · 20VC with Harry Stebbings
McGrath: RLHF and RLVR differ by data quality, not optimization math
“Really, at the end of the day, like, RLHF, RLVR, They're both policy gradient methods, but the, what's different is just like the input data.”
Josh McGrath Dec 31, 2025 ▶ 9:02 [State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI
Zhang: Superhuman computer vision requires RLHF rather than human SFT data
“But if you only do SFT and the SFT data is annotated by human, then your performance is funded by human. You cannot get, kind of, superhuman performance just by, kind of, this kind of data engine approach to use human annotated data and then learn from that. Y…”
Pengchuan Zhang Dec 18, 2025 ▶ 44:33 SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow)
MAD Insight
Łukasz Kaiser: Early RLHF was brittle but crucial for chatbot development
“So it was a bit of a brittle technique, but it was a bit of RL that was extremely crucial to making the models chat.”
Łukasz Kaiser Nov 26, 2025 ▶ 15:53 What’s Next for AI? OpenAI’s Łukasz Kaiser (Transformer Co-Author)
MAD Insight
Lambert: RLVR targets performance characteristics better than traditional RLHF reward models
“These reward models tend to have a lot of problems and you can over optimize them much more easily because the reward models will pick up on features that are maybe emojis or something like this that you don't actually care about where RLVR is much better matc…”
Nathan Lambert Nov 20, 2025 ▶ 1:13:35 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
Emmons: Conversational RLHF rewards undermine automated AI agent workflows
“The other problem I see is the constant reward mechanism of keep the conversation going of it always asks you a follow-up, which is not great for workflows. I want my workflow to stop, not to continuously work.”
Shane Emmons Nov 4, 2025 ▶ 32:01 Building Trustworthy AI: Navigating the Challenges and Future of Agentic Software with Shane Emmons
a16z Assertion Not checkable as stated
Masad: AI models fail to reason on controversial topics due to RLHF
“They can't reason about it because of all the RLHF and all sorts of limitations.”
Amjad Masad Oct 23, 2025 ▶ 48:04 Marc Andreessen & Amjad Masad on “Good Enough” AI, AGI, and the End of Coding
Huyen: Comparative evaluation is significantly easier for humans than absolute scoring
“As humans we tend to, it's very hard to give, like, concrete score. But it's easier to do comparisons, right?”
Chip Huyen Oct 23, 2025 ▶ 16:44 Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix)
SOURCERY Insight
Siddharth: Verifiable domains allow self-play reinforcement learning to replace RLHF
“Now, for these verifiable domains like coding and math, instead of doing reinforcement learning with human feedback, you can do reinforcement learning. Because you can automatically check when you got the correct answer or not in these verifiable domains. And …”
Jonathan Siddharth Oct 10, 2025 ▶ 24:24 Inside The $2.2B AI Research Accelerator | Turing · Sourcery with Molly O'Shea
20VC Assertion Supported
Surge AI is the largest player in RLHF data
“Surge is the largest player in RLHF”
Brendan Foody Sep 15, 2025 ▶ 58:58 Mercor CEO & Co-Founder, Brendan Foody: How They Grew from $1M to $500M in 17 Months · 20VC with Harry Stebbings
Lambert: RLHF can never be permanently solved
“In the same way that chatbot arena can never be saturated. RLHF can never be solved.”
Nathan Lambert Jul 31, 2025 ▶ 16:42 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
20VC Assertion Not checkable as stated
Chen: ChatGPT served as a massive inflection point for Surge AI
“Things definitely hit an excellent point with ChatGPT because I think people just saw how Incredibly valuable human data and RHF was. So definitely chat CPT was an inflection point for us, but even before that we were, we had very strong growth.”
Edwin Chen Jul 21, 2025 ▶ 34:04 Surge CEO & Co-Founder, Edwin Chen: Scaling to $1BN+ in Revenue with NO Funding · 20VC with Harry Stebbings
Diana Hu: XML Prompts Produce Better Output Due to Post-Training RLHF
“We found that it makes it a lot easier for LLMs to follow, because a lot of elements were post-trained in RLHF with kind of XML type of input, and it turns out to produce better results.”
Diana Hu May 30, 2025 ▶ 3:43 State-Of-The-Art Prompting For AI Agents · Y Combinator
Hu: Claude is naturally human-steerable while Llama requires heavy prompting
“One of the things that's known a lot is Claude is sort of the more happy and more human steerable model, and the other one is Lama. Four is one that needs a lot more steering. It's almost like talking to a developer, and part of it could be an artifact of not …”
Diana Hu May 30, 2025 ▶ 26:28 State-Of-The-Art Prompting For AI Agents · Y Combinator
LATENT SPACE Assertion Supported
Brown: GRPO is more memory efficient and easier to distribute
“GRPO is, like, great for, like, leaning heavy on highly parallel inference compute. It's more memory efficient for the actual training process. It's much easier to do in a distributed fashion because you have less gradient syncing and less model weight copies.”
Will Brown May 23, 2025 ▶ 32:07 ⚡️Multi-Turn RL for Multi-Hour Agents — with Will Brown, Prime Intellect
20VC Assertion Supported
Nayak: Scale AI Partnered With OpenAI on Early GPT-2 RLHF
“Scale partnered with open AI very, very early on before chat GPT came out. This was like very early on RLHF when they were trying to tune models to summarize better based off of Reddit passages. And this is on GPT two.”
Aatish Nayak Apr 11, 2025 ▶ 8:22 20Product: How Scale AI and Harvey Build Product | Why PMs Are Wrong: They are not the CEOs of the Product | How to do Pre and Post Mortems Effectively and How to Nail PRDs | The Future of Product Management in a World of AI with Aatish Nayak
Agarwal: Standard RLHF Infrastructure Can Be Repurposed for Model Distillation
“First step is you go to an RLHF, RLXF, whatever framework you have. You turn off the reward term. So you delete the reward part of it. All you are left is some KL term. And now you swap your KL, the anchor policy, to a bigger teacher policy. And there you go. …”
Rishabh Agarwal Mar 23, 2025 ▶ 31:54 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Agarwal: RLHF and LLM distillation can be combined simultaneously
“You can actually combine RLHF and distillation together, because now you're doing two things at the same time.”
Rishabh Agarwal Mar 23, 2025 ▶ 33:17 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Frosst: RLHF data efficiency surprised everyone in AI except OpenAI
“I think that caught pretty much everybody, but the people in OpenAI by surprise was that you can have a relatively small number of examples from people, fine tune the model on that, and then it's a lot easier to work with.”
Nick Frosst Sep 9, 2024 ▶ 4:14 He built Cohere into a $5.5B AI startup; How to Win in AI; & Why LLMs won't lead to AGI. | Nick F... · PMF Show
LATENT SPACE Assertion Supported
Scialom: Meta had to reinvent scaling RLHF without published frontier research
“You have just the basics, but then when it comes to, like, ChatGPT or GPT Instruct or Cloud, No one published the details there. And so we had to reinvent the wheel there in a very short amount of time.”
Thomas Scialom Jul 23, 2024 ▶ 7:53 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
LATENT SPACE Disclosure
Scialom: Smaller Llama 3 models improved via distillation from 405B
“Having bigger models enables to collect better data, for instance, at RLHF stage, because that's the model we use for the annotation. And so we distillate straight forward, like those annotations from this better model to the other models. So I can guarantee y…”
Thomas Scialom Jul 23, 2024 ▶ 14:02 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Scialom: RLHF yields superhuman models because humans judge better than they generate
“And because of that, you can have a model that flats the bad outputs, and learns to only shift towards the best and better and better outputs. And you can even end to superhuman abilities, since that I'm bad at writing a poem, but I'm good at judging which one…”
Thomas Scialom Jul 23, 2024 ▶ 31:56 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Fox: RLHF and dialogue crossed the capability threshold that triggered LLM adoption
“It was the RLHF and the dialogue component of ChatGBT that triggered the takeoff. And that was the capability threshold. Like it was at that time that this capability threshold passed where now it is a prior, like LLMs are a priority for every organization.”
Dylan Fox Apr 8, 2024 ▶ 19:07 This solo founder bet on AI 7 years ago. Now he has 5,000 customers & $115M raised. | Dylan, Foun... · PMF Show
Lambert: AI field hasn't learned much more, just articulates ignorance better
“I think it's, I try to do it every six or 12 months is my current, is my estimated cadence, just to refine the ways that I say things, and people will see That we don't know that much more, but we have a bit of better way of saying what we don't know.”
Nathan Lambert Jan 11, 2024 ▶ 3:58 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Lambert: Major tech companies now realize they need dedicated RLHF teams
“I think any major company that wasn't doing RLHF is now realizing they have to have a team around this.”
Nathan Lambert Jan 11, 2024 ▶ 5:31 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
LATENT SPACE Assertion Not checkable as stated
Lambert: Open-source and research communities lack large-scale RLHF capabilities
“At the same time, we don't have a lot of that in the like open and research communities at the same scale.”
Nathan Lambert Jan 11, 2024 ▶ 5:38 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Lambert: Anthropic is considered the master of RLHF techniques
“Everyone knows Anthropik is kind of the masters of this”
Nathan Lambert Jan 11, 2024 ▶ 5:49 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Lambert: RLHF reward models lack inductive biases for true preferences
“In RLHF calling things a preference model is a little annoying because There's no inductive bias of what a preference is.”
Nathan Lambert Jan 11, 2024 ▶ 12:41 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
LATENT SPACE Assertion Not checkable as stated
Lambert: Frontier AI labs do not recruit economics or social choice academics
“The RLHF techniques that people use were built in, like, labs like OpenAI and DeepMind, where there are some of these people, they have, they, these places do a pretty good job of trying to get these people in the door when you compare them to, like, startups,…”
Nathan Lambert Jan 11, 2024 ▶ 14:06 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
LATENT SPACE Assertion Not checkable as stated
Lambert: Real-Time Human Feedback in LLM RL Loops Is Far From Feasible
“Setting up the infrastructure to take tens of thousands of prompts and generate them and then show them to a human and collect the human responses and then show that Shove that into your training architecture is very far away from working, so we don't really h…”
Nathan Lambert Jan 11, 2024 ▶ 16:13 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Lambert: RLHF operates as a contextual bandit problem, not sequential RL
“What happens is the state is a prompt, and then you do a completion, and then you throw it away, and you grab a new prompt. Where really in like RL, you, as an RL researcher, you would think of this as being like, you take a state, you complete, Get some compl…”
Nathan Lambert Jan 11, 2024 ▶ 16:40 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Lambert: RLHF provides a richer signal per unit of compute than pre-training
“As reinforcement learning is so much less compute, like it, Is a richer signal in terms of its impact, because if they could do what RLHF is doing at pre-training, they would, but they don't know how to have that effect in like a stable manner. Otherwise, ever…”
Nathan Lambert Jan 11, 2024 ▶ 22:14 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Lambert: Instruction tuning is more important than RLHF for most practitioners
“I think for most people, instruction tuning is probably still more important in their day-to-day life. I think instruction tuning works very well. You can write samples by hand that make sense. You can get the model to learn from them. You could do this with v…”
Nathan Lambert Jan 11, 2024 ▶ 28:11 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Lambert: Startups should avoid RLHF unless it offers niche advantage
“I don't really recommend most startups to do it unless it's like going to provide them a clear competitive advantage in their kind of niche. Yeah. Because you're not going to make your model ChatGPT like better than OpenAI or anything like that.”
Nathan Lambert Jan 11, 2024 ▶ 29:37 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Lambert: RLHF fails completely without a KL divergence constraint
“The KL constraint prevents that. There's not that much documented work on that, but there's a lot of people that know if you take that away, it just doesn't work at all.”
Nathan Lambert Jan 11, 2024 ▶ 35:51 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Lambert: RLHF avoids controversial preferences, focusing on correctness and style
“The reason this really is done on a deep level is that you're not actually trying to model any, like, contestable preference in this. Like, you're not trying to go into things that are controversial or anything. It's really, like, the notion of preference is t…”
Nathan Lambert Jan 11, 2024 ▶ 39:02 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Lambert: Annotator disagreement data in RLHF is not used for anything
“I don't really think this disagreement data is used for anything, but it's good to know, like, what the distribution of prompts is, who's doing it, how many samples you have, controlling the workforce.”
Nathan Lambert Jan 11, 2024 ▶ 42:20 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Lambert: A 100x smaller language model filters output better than RLHF
“You could use like a hundred times smaller language model and do much better at filtering than RLHF”
Nathan Lambert Jan 11, 2024 ▶ 45:20 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
LATENT SPACE Assertion Not checkable as stated
Lambert: Zephyr was the first open model to succeed with public RLHF
“I think Zephyr was the first model that showed success with RLHF in the public, but that's a long time from everyone knowing that it was something that people are interested in to having any, like, check mark.”
Nathan Lambert Jan 11, 2024 ▶ 51:42 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Lambert: RLHF Performance Depends on Data and Systems Over RL Details
“It really ends up being kind of, like, gibberish that I think is less important now, because it's more about data and infrastructure than RL details, than, like, value functions and everything. A lot of the papers have different terms in the equations. I think…”
Nathan Lambert Jan 11, 2024 ▶ 57:58 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Lambert: Reinforcement Learning Does Far Less for Alignment Than People Think
“It's like, now you're getting to the point where you don't even really need this to get a good model, so that's why it's like, okay, the RL is such a small part of the actual, like, doing RLHF. Like, RLHF is a metaphor for, like, all language model adaptation,…”
Nathan Lambert Jan 11, 2024 ▶ 58:36 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
LATENT SPACE Assertion Supported
Lambert: RLHF has not been shown to improve underlying model benchmark capabilities
“RLHF is not that shown to improve capabilities yet. I think one of the fun ones is from the GPT-IV technical report. They essentially listed their kind of bogus evaluations, because it's a hilarious table, because it's like LSAT AP exams, and then like AMC-X a…”
Nathan Lambert Jan 11, 2024 ▶ 59:53 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Lambert: Scaling PPO in RLHF is a Nightmare
“PPO is kind of a nightmare to scale.”
Nathan Lambert Jan 11, 2024 ▶ 1:01:41 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
LATENT SPACE Prediction Not checkable as stated
Lambert: Offline RL could take off due to simpler training pipelines
“There's a few papers that people have published Not a lot of traction. I think it could take off. Some people that I know in the RLHF area really think a lot of people are doing this in industry just because it makes the kind of training process simpler and th…”
Nathan Lambert Jan 11, 2024 ▶ 1:04:08 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
LATENT SPACE Assertion Supported
Lambert: Most open-source RLHF training runs only last a few epochs
“Most RLHF is only a few epochs, at least in the open models”
Nathan Lambert Jan 11, 2024 ▶ 1:09:04 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Lambert: DPO is closer to RLHF than RLHF is to RL
“I think DPO is closer to RLHF than RLHF is to RL.”
Nathan Lambert Jan 11, 2024 ▶ 1:12:56 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Lambert: Frontier Labs Lack Visibility into Cross-Model RLHF Sensitivity
“I think big labs are so over-indexed, are indexed on their own base models, so they don't know, like, what's swapping between CloudBase or GPT-IV-Base, how that would change any notion of preference or what you do with RLHF.”
Nathan Lambert Jan 11, 2024 ▶ 1:30:15 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
LATENT SPACE Assertion Supported
Lambert: OpenAI retrains reward models with curated and user prompt mixtures
“And this is like a sort of outer loop optimization that no one in the open is even remotely qualified to talk about, but OpenAI does monitor and they'll like rerun RLHF and train a new reward model with a mixture of their curated data and user prompts to try t…”
Nathan Lambert Jan 11, 2024 ▶ 1:32:27 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
MAD Insight
Kant: Programmatic RL can scale magnitudes larger than human feedback
“And it's that RL loop that is very interesting because since it's programmatic, since we have an Oracle of truth, we can scale this up far larger, right? Magnitudes larger than what you can do with human feedback today.”
Eiso Kant Dec 20, 2023 ▶ 21:27 The Race to Build the Ultimate AI Programmer | Poolside CTO Eiso Kant
MAD Insight
Zhou: Zero-shot LLMs can replace manual human labeling in RLHF workflows
“Which is that why can't it be another LLM or a pipeline of LLMs that can help with that feedback? I think manual labeling is very tedious, especially for our target user, which is a software engineer. And I don't think people should necessarily have to do all …”
Sharon Zhou Nov 8, 2023 ▶ 16:43 Custom LLMs at Scale: Lamini CEO Sharon Zhou’s Playbook for Enterprise AI
a16z Assertion Supported
Dario Amodei says he co-invented RLHF while working at OpenAI
“I was one of the, like, co-inventors of that at OpenAI, but since then it's been, you know, improved to power ChatGPT”
Dario Amodei Sep 25, 2023 ▶ 12:06 Improving AI with Anthropic's Dario Amodei
NO PRIORS Assertion Not checkable as stated
Guo: Major AI Labs Insource Annotators Due to Vendor Quality Deficits
“One thing that I've seen with significant research labs is like still continued insourcing of annotators for both pre-training sets and LHF because some of the external services and marketplaces can't get to the level of quality that they're looking for in par…”
Sarah Guo Sep 14, 2023 ▶ 26:36 No Priors Ep. 32 | With NEAR’s Illia Polosukhin
Reyes: GPT-4's Quality Stems from RLHF, Not Parameter Scale
“GPT-IV is not good because it's large. It's good because it uses reinforcement learning from human feedback, which was discovered in the StrapGPT paper two or three years ago.”
Ricardo Michel Reyes Aug 15, 2023 ▶ 32:19 550 Million People, One Language, Zero Access to Venture Capital
Hotz: RLHF models adopt customer support personalities
“I don't like the RLHF models. I don't like the tuned versions of them. I think that they become, you take on the personality of a customer support agent, right?”
George Hotz Jun 20, 2023 ▶ 1:06:11 Ep 18: Petaflops to the People — with George Hotz of tinycorp
NO PRIORS Insight
Yarats: Rejection sampling significantly boosts LLM quality before full RLHF
“Full blown, like RLHF is, you know, definitely something we're going to look into that, but there is like several many steps that you can have in between that significantly can increase your quality. So for example, I mean like even using something like a reje…”
Denis Yarats Apr 25, 2023 ▶ 20:17 No Priors Ep. 9 | With Perplexity AI’s Aravind Srinivas and Denis Yarats

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.