Mechanistic Interpretability
topic on 4 shows · 14 statements across 9 episodes
Latent Space
Lenny's Podcast
No Priors
the MAD Podcast
14 statements about Mechanistic Interpretability, every show
Kolter: Mechanistic interpretability is not yet a real science
“The problem with McInterp is it's a lot, it's been about sort of testing small hypotheses. Hypothesis. And you know, you have a hypothesis, you'll find some small thing, you'll test that in isolation. But I don't think it's really become a science yet.”
Kolter: Coding agents will revitalize mechanistic interpretability research
“Most fascinating things about coding agents actually is they can do a lot of experimentation in an automated fashion. Yeah. They will give new hope. They'll breathe new life into mechanter research.”
Alex Rives: Mechanistic interpretability will uncover biology inside protein models
“The hope is that you kind of really learn the underlying basis for how it's making the predictions, and so you open up the black box and you can actually understand kind of the biology that the model is representing.”
Understanding model weights and activations is essential for AI safety
“We believe that understanding the internal weights and activations, what is the internal structure, the mathematical structure of these systems is going to be at least part of the solution.”
Smarter AI models will make mechanistic interpretability and tracking more effective
“But as we're starting to have models that are much smarter than us, at least in some important ways, we think that we'll be able to start tracking mechanistic capability much more effectively.”
Kolter: AI agents might turn mechanistic interpretability into a science
“I think that we actually might finally be able to make more what I would consider a science of this through essentially leveraging mass research by agents deployed for this problem.”
Cherny: Anthropic can trace specific neuron activations related to AI deception
“We at this point have like pretty sophisticated technology to understand what's happening in the neurons to trace it. And so for example, like if there's a neuron related to deception, we can start, we're starting to get to the point where we can monitor it an…”
Bissell: Interpretability is rarely applied during training for model design
“Bring interpretability to training, which I don't think has been done all that much before. A lot of this stuff is sort of post-talk poking at models as opposed to actually using this to intentionally design them.”
Interpretability Research Will Explain AI Model Outputs Within Three Years
“I think that if we further that research direction two, three years in the future, we will be able to understand why models say what they'd say.”
Mechanistic Interpretability Scales Without Bottlenecks to Large Models Like DeepSeek
“There's no gap for scale. Like, they've shown that even for the biggest open source models, you like, even like DeepSeq's big models, you, they can do it. And then in general, like, scaling is not the bottleneck.”
Ameisen: Validated circuit models enable predictable steering via feature swapping
“If you understood the circuit well, and if you identified where it's thinking about Huskies or where it's thinking about like kind of like breeding two different breeds, then you should be able to like swap these in and out and get it to kind of like say whate…”
Ameisen: Circuit tracing diagnoses model failures by exposing incorrect internal representations
“Like you, you're not limited to studying what the model can do, right? Like if the model's failing at something like, you know, counting the number of letters in strawberry or whatever you could just try that and try to figure out the circuit for like, well, i…”
Ameisen: Interpretability research has lower entry barriers and low compute needs
“I think for Interp in particular, there's like another thing that makes it easier to transition to, which is maybe two things. One, you can just do it without huge access to compute. Like, there are open source models. You can look at them. A lot of Interp pap…”