Dec 31, 2025 · 21m · latent-space
[State of MechInterp] SAEs in Production, Circuit Tracing, AI4Science, "Pragmatic" Interp — Goodfire
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this NeurIPS interview, Goodfire's Jack and Mark discuss the state of mechanistic interpretability, demonstrating how sparse autoencoders, latent steering, and circuit tracing transition black-box neural networks into practical enterprise solutions and scientific discovery engines.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Jack Merullo forcefully dismisses community readings of Neil Nanda's post as a gross misattribution regarding the viability of interpretability.
Hardest push from the hosts ▶ 17:18 Host clarifies actual industry consensus on managing by outcomesHost Alessio Fanelli directly interjects to reject the guest's framing, clarifying that practitioners are debating outcome-driven steering rather than declaring interpretability dead.
Biggest teaching moment ▶ 18:43 Guest explains Pasteur's Quadrant framework for researchMark Bissell educates the host on Pasteur's Quadrant, breaking down Bohr's pure basic science, Edison's applied focus, and Pasteur's hybrid model.
The host holds their own ▶ 7:11 Host outlines knowledge updating challenges in language modelsAlessio Fanelli demonstrates deep domain familiarity by articulating the specific difficulty of date-conditioned fact editing over unlearning, which the guest affirms with the ROME paper.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Interpreting Diffusion Models with Interactive Concept Painting | 5 | 3 | 1 | 1 | The host asks informed technical questions regarding model unlearning, knowledge updating, and evaluation benchmarks. The guests clarify the distinction between unlearning and suppression, explaining the gradient between rote memorization and reasoning. | |
| Enterprise Production Use Cases and AI for Science | 4 | 2 | 0 | 0 | The host prompts discussion on industrial interpretability use cases and references frontier model cards and AI for science initiatives. The guest shares production case studies, including cost-effective PII scrubbing using feature probes. | |
| Sparse Autoencoders, Circuit Tracing, and Alignment Science | 5 | 4 | 1 | 1 | The host provides a working summary of sparse autoencoders and asks how cross-layer transcoders differ. The guest refines the definition by breaking down cross-layer feature tying and attribution graphs across model representations. | |
| Pragmatic Interpretability and Pasteur's Quadrant Framework | 5 | 4 | 3 | 4 | Discussion centers on pragmatic interpretability; when the guest dismisses external commentary as claiming interpretability is dead, the host pushes back to clarify the nuanced industry critique. The guest then explains Pasteur's Quadrant as a guiding framework. |