RLHF
4 statements across 3 episodes · 3 bullish · 0 bearish · 3 people on the record · first statement Feb 28, 2023 by Raza Habib · across every show →
Everything said about RLHF, oldest first
Feb 28, 2023 positive
Habib: Anthropic matched RLHF performance using AI-generated feedback
“Anthropic had this very exciting paper just a couple of weeks ago where actually we're able to get similar results to RLHF without the H. So just actually having a second model provide the evaluation feedback as well. And that's obviously a lot more scalable.”
May 30, 2025 positive
May 30, 2025 neutral
Hu: Claude is naturally human-steerable while Llama requires heavy prompting
“One of the things that's known a lot is Claude is sort of the more happy and more human steerable model, and the other one is Lama. Four is one that needs a lot more steering. It's almost like talking to a developer, and part of it could be an artifact of not …”
Apr 29, 2026 bullish
Hassabis: Current AI paradigms will be part of final AGI architecture
“The components that you just mentioned, I'm pretty sure will be part of the final architecture for AGI. So I think they've come such a long way now and we've proven out so many things about what they can do. I can't see a world in which we will sort of realize…”