RLHF

5 statements across 5 episodes · 3 bullish · 0 bearish · 5 people on the record · first statement Nov 8, 2023 by Sharon Zhou · across every show →

Everything said about RLHF, oldest first

Nov 8, 2023 positive
Insight
Zhou: Zero-shot LLMs can replace manual human labeling in RLHF workflows
“Which is that why can't it be another LLM or a pipeline of LLMs that can help with that feedback? I think manual labeling is very tedious, especially for our target user, which is a software engineer. And I don't think people should necessarily have to do all …”
Sharon Zhou Nov 8, 2023 ▶ 16:43 Custom LLMs at Scale: Lamini CEO Sharon Zhou’s Playbook for Enterprise AI
Dec 20, 2023 positive
Insight
Kant: Programmatic RL can scale magnitudes larger than human feedback
“And it's that RL loop that is very interesting because since it's programmatic, since we have an Oracle of truth, we can scale this up far larger, right? Magnitudes larger than what you can do with human feedback today.”
Eiso Kant Dec 20, 2023 ▶ 21:27 The Race to Build the Ultimate AI Programmer | Poolside CTO Eiso Kant
Nov 20, 2025 positive
Insight
Lambert: RLVR targets performance characteristics better than traditional RLHF reward models
“These reward models tend to have a lot of problems and you can over optimize them much more easily because the reward models will pick up on features that are maybe emojis or something like this that you don't actually care about where RLVR is much better matc…”
Nathan Lambert Nov 20, 2025 ▶ 1:13:35 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
Nov 26, 2025 neutral
Insight
Łukasz Kaiser: Early RLHF was brittle but crucial for chatbot development
“So it was a bit of a brittle technique, but it was a bit of RL that was extremely crucial to making the models chat.”
Łukasz Kaiser Nov 26, 2025 ▶ 15:53 What’s Next for AI? OpenAI’s Łukasz Kaiser (Transformer Co-Author)
Aug 6, 2026 neutral
Assertion Supported
Wolf: Frontier AI Training Has Shifted From RLHF to Pure RL
“What we know though, is we moved from this pure, like human data, you know, that was first just pre-training on human data and then also aligning with like human preferences that was called RLHF, where we had a lot of human in the loop and human data. To like …”
Thomas Wolf Aug 6, 2026 ▶ 29:08 “OpenAI’s Model Hacked Us” - Hugging Face’s Thomas Wolf
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.