Using chain-of-thought as an RL reward destroys model interpretability
Julian Schrittwieser · Are We Misreading the AI Exponential? Julian Schrittwieser on Move 37 & Scaling RL (Anthropic) · Oct 23, 2025 · at 58:25
Julian Schrittwieser, AI researcher at Anthropic, discusses the trade-offs between reinforcement learning (RL) training techniques and maintaining visibility into AI internal reasoning.
“If you're not careful with RL, you can make interpretability harder. For example, one Common thing with modern models is they do reasoning with the chain of thought. You could look at the chain of thoughts to, you know, see what are the model internal thoughts, and then you could also have a thought that, oh, maybe I should use that as a reward signal in RL and punish the model if it thinks the wrong thing, but then suddenly you completely destroyed your interpretability angle, so you sort of have to be careful that, yeah, you don't Do RL on the signals that you actually want to use to interpret what the model is thinking of doing.”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →