mechanistic interpretability

9 statements across 5 episodes · 7 bullish · 1 bearish · 5 people on the record · first statement Dec 23, 2024 by Luca Soldani · across every show →

Everything said about mechanistic interpretability, oldest first

Dec 23, 2024 positive
Opinion
Soldani: Mechanistic interpretability research is impossible without open models
“There is a large swath of research on modeling, on how these models behave, on evaluation, on inference, on mechanistic interpretability that could not happen at all. If you didn't have open models.”
Luca Soldani Dec 23, 2024 ▶ 3:08 Best of 2024: Open Models [LS LIVE! at NeurIPS 2024]
Jun 6, 2025 positive
Insight
Ameisen: Circuit tracing diagnoses model failures by exposing incorrect internal representations
“Like you, you're not limited to studying what the model can do, right? Like if the model's failing at something like, you know, counting the number of letters in strawberry or whatever you could just try that and try to figure out the circuit for like, well, i…”
Emmanuel Ameisen Jun 6, 2025 ▶ 23:02 The Utility of Interpretability — Emmanuel Amiesen
Jun 6, 2025 positive
Insight
Ameisen: Interpretability research has lower entry barriers and low compute needs
“I think for Interp in particular, there's like another thing that makes it easier to transition to, which is maybe two things. One, you can just do it without huge access to compute. Like, there are open source models. You can look at them. A lot of Interp pap…”
Emmanuel Ameisen Jun 6, 2025 ▶ 28:56 The Utility of Interpretability — Emmanuel Amiesen
Jun 6, 2025 positive
Insight
Ameisen: Validated circuit models enable predictable steering via feature swapping
“If you understood the circuit well, and if you identified where it's thinking about Huskies or where it's thinking about like kind of like breeding two different breeds, then you should be able to like swap these in and out and get it to kind of like say whate…”
Emmanuel Ameisen Jun 6, 2025 ▶ 20:45 The Utility of Interpretability — Emmanuel Amiesen
Nov 14, 2025 bullish
Prediction Not checkable as stated
Interpretability Research Will Explain AI Model Outputs Within Three Years
“I think that if we further that research direction two, three years in the future, we will be able to understand why models say what they'd say.”
Deedy Das Nov 14, 2025 ▶ 51:27 Anthropic, Glean & OpenRouter: How AI Moats Are Built with Deedy Das of Menlo Ventures
Nov 14, 2025 positive
Assertion Supported
Mechanistic Interpretability Scales Without Bottlenecks to Large Models Like DeepSeek
“There's no gap for scale. Like, they've shown that even for the biggest open source models, you like, even like DeepSeq's big models, you, they can do it. And then in general, like, scaling is not the bottleneck.”
Deedy Das Nov 14, 2025 ▶ 52:11 Anthropic, Glean & OpenRouter: How AI Moats Are Built with Deedy Das of Menlo Ventures
Feb 5, 2026
Insight
Bissell: Interpretability is rarely applied during training for model design
“Bring interpretability to training, which I don't think has been done all that much before. A lot of this stuff is sort of post-talk poking at models as opposed to actually using this to intentionally design them.”
Mark Bissell Feb 5, 2026 ▶ 7:58 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Jun 22, 2026 bullish
Prediction Not checkable as stated
Kolter: Coding agents will revitalize mechanistic interpretability research
“Most fascinating things about coding agents actually is they can do a lot of experimentation in an automated fashion. Yeah. They will give new hope. They'll breathe new life into mechanter research.”
Zico Kolter Jun 22, 2026 ▶ 17:58 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Jun 22, 2026 negative
Opinion
Kolter: Mechanistic interpretability is not yet a real science
“The problem with McInterp is it's a lot, it's been about sort of testing small hypotheses. Hypothesis. And you know, you have a hypothesis, you'll find some small thing, you'll test that in isolation. But I don't think it's really become a science yet.”
Zico Kolter Jun 22, 2026 ▶ 17:29 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.