Insight certainty 4/5 debate potential 3/5

Using chain-of-thought as an RL reward destroys model interpretability

Julian Schrittwieser · Are We Misreading the AI Exponential? Julian Schrittwieser on Move 37 & Scaling RL (Anthropic) · Oct 23, 2025 · at 58:25

Julian Schrittwieser, AI researcher at Anthropic, discusses the trade-offs between reinforcement learning (RL) training techniques and maintaining visibility into AI internal reasoning.

0:00 / 0:34exact quote · 34.3s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“If you're not careful with RL, you can make interpretability harder. For example, one Common thing with modern models is they do reasoning with the chain of thought. You could look at the chain of thoughts to, you know, see what are the model internal thoughts, and then you could also have a thought that, oh, maybe I should use that as a reward signal in RL and punish the model if it thinks the wrong thing, but then suddenly you completely destroyed your interpretability angle, so you sort of have to be careful that, yeah, you don't Do RL on the signals that you actually want to use to interpret what the model is thinking of doing.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Julian Schrittwieser

Prediction Not checkable as stated
AI will make Nobel Prize-level scientific discoveries by 2027 or 2028
“I think my guess for that level of capability might be maybe 2027. I think we're probably not going to find out for quite some time afterwards because of the delay in getting prices. But I think by 20, 27, 20, 28, I think extremely likely that the models will …”
Julian Schrittwieser Oct 23, 2025 ▶ 15:35 Are We Misreading the AI Exponential? Julian Schrittwieser on Move 37 & Scaling RL (Anthropic)
Prediction Held up
Top AI models will work autonomously for full days within two years
“In a year from now, maybe two years from now, it's the top models are going to be able to work completely on their own for like a whole day or more”
Julian Schrittwieser Oct 23, 2025 ▶ 3:01 Are We Misreading the AI Exponential? Julian Schrittwieser on Move 37 & Scaling RL (Anthropic)
Opinion
Valuations for OpenAI, Anthropic, and Google are fairly conservative
“If you look at OpenAI, if you look at Anthropic, if you look at Google, those evaluations, those revenue numbers are actually fairly conservative.”
Julian Schrittwieser Oct 23, 2025 ▶ 3:32 Are We Misreading the AI Exponential? Julian Schrittwieser on Move 37 & Scaling RL (Anthropic)
Prediction Not checkable as stated
A sudden AI singularity or intelligence explosion is extremely unlikely
“Yeah, I think a true discontinuity is extremely unlikely from, you know, obviously AI researchers are already using AI to accelerate themselves. And so what's, what's already happening and like what is likely to continue to happening is that we see like a smoo…”
Julian Schrittwieser Oct 23, 2025 ▶ 17:11 Are We Misreading the AI Exponential? Julian Schrittwieser on Move 37 & Scaling RL (Anthropic)
Assertion Not checkable as stated
AI autonomous task duration doubles every three to four months
“We are seeing this very consistent improvement over many, many years where every say like, you know, three, four months is able to like do a task that is twice as long as before completely on its own.”
Julian Schrittwieser Oct 23, 2025 ▶ 2:41 Are We Misreading the AI Exponential? Julian Schrittwieser on Move 37 & Scaling RL (Anthropic)
Opinion
Schrittwieser: Wider AI ecosystem may face bubble while frontier labs thrive
“There may simultaneously be like some sort of bubble in, you know, the wider ecosystem, while at the same time, the frontier labs on a very solid trajectory, having a lot of revenue, making a lot of money.”
Julian Schrittwieser Oct 23, 2025 ▶ 4:13 Are We Misreading the AI Exponential? Julian Schrittwieser on Move 37 & Scaling RL (Anthropic)
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.