White: RLHF fails on scientific hypotheses by ignoring impact and information gain
Andrew White · 🔬 From Red Teaming GPT-4 to Automating Drug Discovery: The Future of AI in Science — Andrew White · Jan 28, 2026 · at 20:35
Andrew White (Future House) discusses experiments using human feedback to train AI agents to generate scientific hypotheses.
“We learned a lot about how bad our LHF is with people, just like people pay really attention to the tone, to the details, to like how many specific facts or figures on the hypothesis, right? Like actionability about like if the experiment is feasible, but what people didn't really pay attention to is like, I don't know how to describe this, but like if this hypothesis is true, how does it change the world? If the hypothesis is false, how does it change the world? This like how much information do you gain? It's not really information, but like impact or something. And that really didn't come through from those things.”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →