Dec 31, 2025 · 21m · latent-space

[State of MechInterp] SAEs in Production, Circuit Tracing, AI4Science, "Pragmatic" Interp — Goodfire

Mark Bissell · 9m spoken Jack Merullo · 7m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this NeurIPS interview, Goodfire's Jack and Mark discuss the state of mechanistic interpretability, demonstrating how sparse autoencoders, latent steering, and circuit tracing transition black-box neural networks into practical enterprise solutions and scientific discovery engines.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 4.8 Guest teaching 3.3 Guest disagreement 1.3 The hosts pushing back 1.5
05100:0010:0020:002:56–8:12 · The hosts as informed peer 5/10 Interpreting Diffusion Models with Interactive Concept Painting The host asks informed technical questions regarding model unlearning, knowledge updating, and evaluation benchmarks. The guests clarify the distinction between unlearning and suppression, explaining the gradient between rote memorization and reasoning.8:13–11:17 · The hosts as informed peer 4/10 Enterprise Production Use Cases and AI for Science The host prompts discussion on industrial interpretability use cases and references frontier model cards and AI for science initiatives. The guest shares production case studies, including cost-effective PII scrubbing using feature probes.11:17–15:43 · The hosts as informed peer 5/10 Sparse Autoencoders, Circuit Tracing, and Alignment Science The host provides a working summary of sparse autoencoders and asks how cross-layer transcoders differ. The guest refines the definition by breaking down cross-layer feature tying and attribution graphs across model representations.15:44–20:26 · The hosts as informed peer 5/10 Pragmatic Interpretability and Pasteur's Quadrant Framework Discussion centers on pragmatic interpretability; when the guest dismisses external commentary as claiming interpretability is dead, the host pushes back to clarify the nuanced industry critique. The guest then explains Pasteur's Quadrant as a guiding framework.2:56–8:12 · Guest teaching 3/10 Interpreting Diffusion Models with Interactive Concept Painting The host asks informed technical questions regarding model unlearning, knowledge updating, and evaluation benchmarks. The guests clarify the distinction between unlearning and suppression, explaining the gradient between rote memorization and reasoning.8:13–11:17 · Guest teaching 2/10 Enterprise Production Use Cases and AI for Science The host prompts discussion on industrial interpretability use cases and references frontier model cards and AI for science initiatives. The guest shares production case studies, including cost-effective PII scrubbing using feature probes.11:17–15:43 · Guest teaching 4/10 Sparse Autoencoders, Circuit Tracing, and Alignment Science The host provides a working summary of sparse autoencoders and asks how cross-layer transcoders differ. The guest refines the definition by breaking down cross-layer feature tying and attribution graphs across model representations.15:44–20:26 · Guest teaching 4/10 Pragmatic Interpretability and Pasteur's Quadrant Framework Discussion centers on pragmatic interpretability; when the guest dismisses external commentary as claiming interpretability is dead, the host pushes back to clarify the nuanced industry critique. The guest then explains Pasteur's Quadrant as a guiding framework.2:56–8:12 · Guest disagreement 1/10 Interpreting Diffusion Models with Interactive Concept Painting The host asks informed technical questions regarding model unlearning, knowledge updating, and evaluation benchmarks. The guests clarify the distinction between unlearning and suppression, explaining the gradient between rote memorization and reasoning.8:13–11:17 · Guest disagreement 0/10 Enterprise Production Use Cases and AI for Science The host prompts discussion on industrial interpretability use cases and references frontier model cards and AI for science initiatives. The guest shares production case studies, including cost-effective PII scrubbing using feature probes.11:17–15:43 · Guest disagreement 1/10 Sparse Autoencoders, Circuit Tracing, and Alignment Science The host provides a working summary of sparse autoencoders and asks how cross-layer transcoders differ. The guest refines the definition by breaking down cross-layer feature tying and attribution graphs across model representations.15:44–20:26 · Guest disagreement 3/10 Pragmatic Interpretability and Pasteur's Quadrant Framework Discussion centers on pragmatic interpretability; when the guest dismisses external commentary as claiming interpretability is dead, the host pushes back to clarify the nuanced industry critique. The guest then explains Pasteur's Quadrant as a guiding framework.2:56–8:12 · The hosts pushing back 1/10 Interpreting Diffusion Models with Interactive Concept Painting The host asks informed technical questions regarding model unlearning, knowledge updating, and evaluation benchmarks. The guests clarify the distinction between unlearning and suppression, explaining the gradient between rote memorization and reasoning.8:13–11:17 · The hosts pushing back 0/10 Enterprise Production Use Cases and AI for Science The host prompts discussion on industrial interpretability use cases and references frontier model cards and AI for science initiatives. The guest shares production case studies, including cost-effective PII scrubbing using feature probes.11:17–15:43 · The hosts pushing back 1/10 Sparse Autoencoders, Circuit Tracing, and Alignment Science The host provides a working summary of sparse autoencoders and asks how cross-layer transcoders differ. The guest refines the definition by breaking down cross-layer feature tying and attribution graphs across model representations.15:44–20:26 · The hosts pushing back 4/10 Pragmatic Interpretability and Pasteur's Quadrant Framework Discussion centers on pragmatic interpretability; when the guest dismisses external commentary as claiming interpretability is dead, the host pushes back to clarify the nuanced industry critique. The guest then explains Pasteur's Quadrant as a guiding framework.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 17:06 Guest rejects pessimistic interpretations of pragmatic interpretability

Jack Merullo forcefully dismisses community readings of Neil Nanda's post as a gross misattribution regarding the viability of interpretability.

Hardest push from the hosts ▶ 17:18 Host clarifies actual industry consensus on managing by outcomes

Host Alessio Fanelli directly interjects to reject the guest's framing, clarifying that practitioners are debating outcome-driven steering rather than declaring interpretability dead.

Biggest teaching moment ▶ 18:43 Guest explains Pasteur's Quadrant framework for research

Mark Bissell educates the host on Pasteur's Quadrant, breaking down Bohr's pure basic science, Edison's applied focus, and Pasteur's hybrid model.

The host holds their own ▶ 7:11 Host outlines knowledge updating challenges in language models

Alessio Fanelli demonstrates deep domain familiarity by articulating the specific difficulty of date-conditioned fact editing over unlearning, which the guest affirms with the ROME paper.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Interpreting Diffusion Models with Interactive Concept Painting 5311 The host asks informed technical questions regarding model unlearning, knowledge updating, and evaluation benchmarks. The guests clarify the distinction between unlearning and suppression, explaining the gradient between rote memorization and reasoning.
Enterprise Production Use Cases and AI for Science 4200 The host prompts discussion on industrial interpretability use cases and references frontier model cards and AI for science initiatives. The guest shares production case studies, including cost-effective PII scrubbing using feature probes.
Sparse Autoencoders, Circuit Tracing, and Alignment Science 5411 The host provides a working summary of sparse autoencoders and asks how cross-layer transcoders differ. The guest refines the definition by breaking down cross-layer feature tying and attribution graphs across model representations.
Pragmatic Interpretability and Pasteur's Quadrant Framework 5434 Discussion centers on pragmatic interpretability; when the guest dismisses external commentary as claiming interpretability is dead, the host pushes back to clarify the nuanced industry critique. The guest then explains Pasteur's Quadrant as a guiding framework.

Statements from this episode (6)

Insight
Bissell: Mechanistic interpretability provides power-user tools for manipulating AI models
“Interpretability gives you a set of, I think of it almost as like power user tools for accessing models and doing things with them that you might not have realized you could.”
Mark Bissell Dec 31, 2025 ▶ 3:26
Assertion Supported
Bissell: Interpretability allows direct painting into a diffusion model's mental map
“Using interpretability techniques, you can sort of like plug directly into the mind of the model, and you get a two D canvas where you can basically like paint directly into its mental map of the image. And so we used unsupervised techniques to basically figur…”
Mark Bissell Dec 31, 2025 ▶ 3:47
Insight
Merullo: LLM memorization spans a gradient from reasoning to rote recall
“You can actually see, like the way that we, like, disentangle memorization, you can kind of see this like, gradient of memorization in between both mechanistically and behaviorally with, like, logical reasoning tasks being quite distinct from rote memorization…”
Jack Merullo Dec 31, 2025 ▶ 6:02
Opinion
Merullo: Current machine unlearning techniques merely suppress data rather than removing it
“I would describe it more as not unlearning, but maybe suppression. I think there's, like, really, like, I guess, guarantees that you've fully removed information from a model is, is, I don't think it's been convincingly showed anywhere yet”
Jack Merullo Dec 31, 2025 ▶ 6:44
Assertion Supported
Bissell: Rakuten uses interpretability in production to scrub PII from customer chats
“From Goodfire's perspective, you know, we, so like one of our partners, Rakuten, is deploying an interpretability based tool in production with one of their language agents. This is a really cool use case where if you, What they needed to do was take chats bet…”
Mark Bissell Dec 31, 2025 ▶ 8:59
Insight
Bissell: Probing internal model features matches LLM-as-a-judge quality at 500x lower cost
“If you ask that model, try to like use it as an LLM as a judge, it's not very good. But if you probe its mind and you sort of detect when the features related to personally identifiable information are firing, that gets you the highest recall of anything. It's…”
Mark Bissell Dec 31, 2025 ▶ 9:53
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.