Mechanistic Interpretability

topic on 4 shows · 14 statements across 9 episodes

Latent Space Lenny's Podcast No Priors the MAD Podcast

14 statements about Mechanistic Interpretability, every show

Kolter: Mechanistic interpretability is not yet a real science
“The problem with McInterp is it's a lot, it's been about sort of testing small hypotheses. Hypothesis. And you know, you have a hypothesis, you'll find some small thing, you'll test that in isolation. But I don't think it's really become a science yet.”
Zico Kolter Jun 22, 2026 ▶ 17:29 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
LATENT SPACE Prediction Not checkable as stated
Kolter: Coding agents will revitalize mechanistic interpretability research
“Most fascinating things about coding agents actually is they can do a lot of experimentation in an automated fashion. Yeah. They will give new hope. They'll breathe new life into mechanter research.”
Zico Kolter Jun 22, 2026 ▶ 17:58 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
NO PRIORS Prediction Not checkable as stated
Alex Rives: Mechanistic interpretability will uncover biology inside protein models
“The hope is that you kind of really learn the underlying basis for how it's making the predictions, and so you open up the black box and you can actually understand kind of the biology that the model is representing.”
Alex Rives Jun 10, 2026 ▶ 16:49 “Curing All Disease by next century is too conservative" - Mark Zuckerberg
NO PRIORS Opinion
Understanding model weights and activations is essential for AI safety
“We believe that understanding the internal weights and activations, what is the internal structure, the mathematical structure of these systems is going to be at least part of the solution.”
Maxim Bar Kogan May 28, 2026 ▶ 21:49 Building an AI Guardian for Enterprise with Onyx Security CEO Maxim Bar Kogan
NO PRIORS Prediction Not checkable as stated
Smarter AI models will make mechanistic interpretability and tracking more effective
“But as we're starting to have models that are much smarter than us, at least in some important ways, we think that we'll be able to start tracking mechanistic capability much more effectively.”
Maxim Bar Kogan May 28, 2026 ▶ 22:44 Building an AI Guardian for Enterprise with Onyx Security CEO Maxim Bar Kogan
MAD Prediction Not checkable as stated
Kolter: AI agents might turn mechanistic interpretability into a science
“I think that we actually might finally be able to make more what I would consider a science of this through essentially leveraging mass research by agents deployed for this problem.”
Zico Kolter May 7, 2026 ▶ 1:02:13 OpenAI Board Member Zico Kolter: Modern AI Is Just 200 Lines of Code
LENNY'S PODCAST Assertion Supported
Cherny: Anthropic can trace specific neuron activations related to AI deception
“We at this point have like pretty sophisticated technology to understand what's happening in the neurons to trace it. And so for example, like if there's a neuron related to deception, we can start, we're starting to get to the point where we can monitor it an…”
Boris Cherny Feb 19, 2026 ▶ 54:46 Head of Claude Code: What happens after coding is solved | Boris Cherny
Bissell: Interpretability is rarely applied during training for model design
“Bring interpretability to training, which I don't think has been done all that much before. A lot of this stuff is sort of post-talk poking at models as opposed to actually using this to intentionally design them.”
Mark Bissell Feb 5, 2026 ▶ 7:58 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
LATENT SPACE Prediction Not checkable as stated
Interpretability Research Will Explain AI Model Outputs Within Three Years
“I think that if we further that research direction two, three years in the future, we will be able to understand why models say what they'd say.”
Deedy Das Nov 14, 2025 ▶ 51:27 Anthropic, Glean & OpenRouter: How AI Moats Are Built with Deedy Das of Menlo Ventures
LATENT SPACE Assertion Supported
Mechanistic Interpretability Scales Without Bottlenecks to Large Models Like DeepSeek
“There's no gap for scale. Like, they've shown that even for the biggest open source models, you like, even like DeepSeq's big models, you, they can do it. And then in general, like, scaling is not the bottleneck.”
Deedy Das Nov 14, 2025 ▶ 52:11 Anthropic, Glean & OpenRouter: How AI Moats Are Built with Deedy Das of Menlo Ventures
Ameisen: Validated circuit models enable predictable steering via feature swapping
“If you understood the circuit well, and if you identified where it's thinking about Huskies or where it's thinking about like kind of like breeding two different breeds, then you should be able to like swap these in and out and get it to kind of like say whate…”
Emmanuel Ameisen Jun 6, 2025 ▶ 20:45 The Utility of Interpretability — Emmanuel Amiesen
Ameisen: Circuit tracing diagnoses model failures by exposing incorrect internal representations
“Like you, you're not limited to studying what the model can do, right? Like if the model's failing at something like, you know, counting the number of letters in strawberry or whatever you could just try that and try to figure out the circuit for like, well, i…”
Emmanuel Ameisen Jun 6, 2025 ▶ 23:02 The Utility of Interpretability — Emmanuel Amiesen
Ameisen: Interpretability research has lower entry barriers and low compute needs
“I think for Interp in particular, there's like another thing that makes it easier to transition to, which is maybe two things. One, you can just do it without huge access to compute. Like, there are open source models. You can look at them. A lot of Interp pap…”
Emmanuel Ameisen Jun 6, 2025 ▶ 28:56 The Utility of Interpretability — Emmanuel Amiesen
Soldani: Mechanistic interpretability research is impossible without open models
“There is a large swath of research on modeling, on how these models behave, on evaluation, on inference, on mechanistic interpretability that could not happen at all. If you didn't have open models.”
Luca Soldani Dec 23, 2024 ▶ 3:08 Best of 2024: Open Models [LS LIVE! at NeurIPS 2024]

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.