White: Writing bulletproof RL verifiers is far harder than supervised training
Andrew White · 🔬 From Red Teaming GPT-4 to Automating Drug Discovery: The Future of AI in Science — Andrew White · Jan 28, 2026 · at 1:12:03
AI researcher Andrew White discusses the EtherZero project, which applied reinforcement learning and verifiable rewards to chemistry and molecular design.
“Pre-training or training transformers, you know on just data, like just supervised training where you just have the inputs and the outputs directly, very nice, relaxing, you know, like things are always robust, you know, things go pretty smoothly. When we do these verifiable rewards where you have to like write a bulletproof verifier, it is really difficult. And we had so many models trained only to find out they were hacking some other like random thing in our setup. It's really hard.”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →