Interpretability
topic on 5 shows · 14 statements across 12 episodes
the Y Combinator Startup Podcast
Latent Space
No Priors
WTF is with Nikhil Kamath
the MAD Podcast
14 statements about Interpretability, every show
Becker: METR Uses Black-Box Methods Over Interpretability for AI Monitoring
“Usually this is black box, not, not white box in, in, in my understanding in, in current work. So, so not using interpretability, but you can imagine in principle doing, doing, doing something more white box.”
Amodei: Interpretability research has identified concept neurons and rhyming circuits in LLMs
“We've been able to find, you know, neurons that correspond to very specific concepts, neural circuits that correspond to, you know, keep track of how to do rhymes in poetry, and so we're starting to understand what these models Do, right?”
Deng: Interpretability Will Unlock the Next Frontier of AI Models
“We really believe that interpretability will unlock the new generation, next frontier of safe and powerful AI models.”
Deng: Computational Neuroscientists Are Moving to AI Interpretability for Unfettered Experimental Access
“When we talk to a lot of computational neuroscientists, they Moved to enter because they were like, look, we have unfettered access to this artificial, intelligent mind. It's so much, you have access to everything. You can run as many ablations and experiments…”
Bissell: Mechanistic interpretability provides power-user tools for manipulating AI models
“Interpretability gives you a set of, I think of it almost as like power user tools for accessing models and doing things with them that you might not have realized you could.”
Rumbelow: Interpretability turns neural networks into scientific discovery tools
“If you've got really good interp, you can start to reframe neural networks, not as just a tool for automating things that we already know how to do, but as a tool for discovery, as like a lens through which you can see patterns in data that would otherwise
Be …”
Using chain-of-thought as an RL reward destroys model interpretability
“If you're not careful with RL, you can make interpretability harder. For example, one Common thing with modern models is they do reasoning with the chain of thought. You could look at the chain of thoughts to, you know, see what are the model internal thoughts…”
Krishnan: Undetectable Backdoors in Coding Assistants Threaten Critical Infrastructure
“Interpretability is still a nascent field and you could very easily see ways where You plug a model into a cursor or windsurf and you generate a piece of code. And then two years down the road, it turns out that code had a little if statement saying, if I'm ru…”
Kaplan: Interpretability is like neuroscience but with complete observability
“I would say that interpretability is a lot more like biology. It's a lot more like neuroscience. So I think those are kind of the tools. There is some more, more, more mathematics there, but I think it's more like trying to understand the features of the brain…”
Ameisen: Anthropic's Open Tool Traces Internal States in Gemma 2 2B
“And then the release this week sort of lets anyone do it for a set of open source models. So notably maybe the most easy one here is like Gemma two to be. So you can sort of like think of some prompt and you kind of like can explain any like token that the mod…”
Ameisen: Anthropic publishes interpretability research to recruit more researchers
“The reason for publishing this is that we think interpretably is important. We think it's tractable, and we think more people should work on it. And so publishing it helps us like accomplish with these goals all these goals, which we think are just like crucia…”
Bach: AI Model Interpretability Requires Automated Reverse-Engineering Systems
“In a way these models are implemented in operator language in which they are performing certain things. But the operator language itself is so complex that it's no longer readable in a way. It goes beyond what you could engineer by hand or what you can reverse…”
Uszkoreit: Interpretable theories for complex deep learning systems exceed human cognitive limits
“There are people trying, and I think it's worth trying. I, I'm not super optimistic about that. I think it'll work for some cases, right, where it's simple enough that we can get it. I think there are many cases where it just isn't, right? Like, say, climate a…”
Koller: AI requires synthesizing deep learning with causal and interpretable models
“What I think we're starting to see right now is a the pendulum starting to swing back in the sense that there is a greater understanding that you really need a bit of both. You need that hugely powerful pattern recognition that we get from deep learning, but y…”