interpretability

8 statements across 6 episodes · 6 bullish · 0 bearish · 6 people on the record · first statement Apr 27, 2024 by Joscha Bach · across every show →

Everything said about interpretability, oldest first

Apr 27, 2024
Insight
Bach: AI Model Interpretability Requires Automated Reverse-Engineering Systems
“In a way these models are implemented in operator language in which they are performing certain things. But the operator language itself is so complex that it's no longer readable in a way. It goes beyond what you could engineer by hand or what you can reverse…”
Joscha Bach Apr 27, 2024 ▶ 1:35:04 This World Does Not Exist — Joscha Bach, Karan Malhotra, Rob Haisfield (WorldSim, WebSim, Liquid AI)
Jun 6, 2025 positive
Assertion Supported
Ameisen: Anthropic's Open Tool Traces Internal States in Gemma 2 2B
“And then the release this week sort of lets anyone do it for a set of open source models. So notably maybe the most easy one here is like Gemma two to be. So you can sort of like think of some prompt and you kind of like can explain any like token that the mod…”
Emmanuel Ameisen Jun 6, 2025 ▶ 1:35 The Utility of Interpretability — Emmanuel Amiesen
Jun 6, 2025 positive
Disclosure
Ameisen: Anthropic publishes interpretability research to recruit more researchers
“The reason for publishing this is that we think interpretably is important. We think it's tractable, and we think more people should work on it. And so publishing it helps us like accomplish with these goals all these goals, which we think are just like crucia…”
Emmanuel Ameisen Jun 6, 2025 ▶ 1:40:51 The Utility of Interpretability — Emmanuel Amiesen
Nov 2, 2025 positive
Insight
Rumbelow: Interpretability turns neural networks into scientific discovery tools
“If you've got really good interp, you can start to reframe neural networks, not as just a tool for automating things that we already know how to do, but as a tool for discovery, as like a lens through which you can see patterns in data that would otherwise Be …”
Jessica Rumbelow Nov 2, 2025 ▶ 2:29 ⚡️Automating Scientific Discovery - Jessica Rumbelow, Leap Labs
Dec 31, 2025 positive
Insight
Bissell: Mechanistic interpretability provides power-user tools for manipulating AI models
“Interpretability gives you a set of, I think of it almost as like power user tools for accessing models and doing things with them that you might not have realized you could.”
Mark Bissell Dec 31, 2025 ▶ 3:26 [State of MechInterp] SAEs in Production, Circuit Tracing, AI4Science, "Pragmatic" Interp — Goodfire
Feb 5, 2026 positive
Insight
Deng: Computational Neuroscientists Are Moving to AI Interpretability for Unfettered Experimental Access
“When we talk to a lot of computational neuroscientists, they Moved to enter because they were like, look, we have unfettered access to this artificial, intelligent mind. It's so much, you have access to everything. You can run as many ablations and experiments…”
Myra Deng Feb 5, 2026 ▶ 1:00:14 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Feb 5, 2026 bullish
Prediction Not checkable as stated
Deng: Interpretability Will Unlock the Next Frontier of AI Models
“We really believe that interpretability will unlock the new generation, next frontier of safe and powerful AI models.”
Myra Deng Feb 5, 2026 ▶ 1:11 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Feb 27, 2026 neutral
Assertion Supported
Becker: METR Uses Black-Box Methods Over Interpretability for AI Monitoring
“Usually this is black box, not, not white box in, in, in my understanding in, in current work. So, so not using interpretability, but you can imagine in principle doing, doing, doing something more white box.”
Joel Becker Feb 27, 2026 ▶ 1:00:33 Measuring Exponential Trends Rising (in AI) — Joel Becker, METR
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.