Jun 6, 2025 · 1h 53m · latent-space
The Utility of Interpretability — Emmanuel Amiesen
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
Anthropic research scientist Emmanuel Ameisen discusses the open-source release of circuit tracing tools for Gemma 2B, demonstrating how sparse autoencoders and attribution graphs expose internal reasoning, backward planning, and AI safety mechanisms in large language models.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Emmanuel explicitly refutes the host's framing that research is inherently more valuable than engineering, arguing that rapid execution and systems engineering make up 90% of real breakthrough value.
Hardest push from the hosts ▶ 13:51 Host presses on hidden limitations of circuit visualizationsThe host directly confronts Emmanuel on whether the Neuronpedia circuit graphs look 'too easy' and 'too clean,' demanding to know what limitations and skeletons are being obscured.
Biggest teaching moment ▶ 1:31:03 Emmanuel exposes unfaithful chain-of-thought and deceptive reasoningEmmanuel demonstrates step-by-step how Claude reverse-engineers a false mathematical output to align with a user prompt hint, proving that surface chains of thought can mask deceptive inner computation.
The host holds their own ▶ 1:10:13 Host introduces Sapir-Whorf linguistic framework into feature sharingThe host synthesizes Emmanuel's cross-lingual feature findings with the Sapir-Whorf linguistic hypothesis and Google's empirical research on massive-context in-context translation.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Exploring Open Questions in Circuit Tracing and Model Reasoning | 4 | 5 | 0 | 0 | Vibhu asks open-ended exploratory questions about how practitioners can leverage the newly open-sourced circuit tracing library. Emmanuel details three tiers of research engagement, ranging from probing multi-hop reasoning in Gemma to extending the core graph generation methods. | |
| Live Demonstration of Circuit Tracing on Neuronpedia | 4 | 6 | 0 | 0 | Emmanuel shares his screen and conducts a live walk-through on Neuronpedia using a podcast completion prompt. He explains how intermediate features activate, how to trace attributions backward, and how causal interventions can be run directly on Colab. | |
| Unpacking Superposition, Errors, and Limitations in Circuit Graphs | 5 | 7 | 1 | 5 | The host pushes back on the clean visualization, asking where the 'skeletons' and superposition are hidden. Emmanuel transparently breaks down graph error diamonds and explains that attention layers are currently left unmodeled. | |
| Interactive Feature Probing: Pomsky Demonstration and Failure Analysis | 5 | 4 | 0 | 0 | Vibhu shares his screen to demonstrate a real-time Pomsky dog breed experiment using Neuronpedia. Emmanuel provides suggestions on swapping breed features to causally verify circuit hypotheses. | |
| Studio Introduction and Transitioning from ML Engineering to Research | 4 | 6 | 3 | 3 | When the host asserts that research is significantly more valuable than engineering, Emmanuel pushes back. He argues that research ideas are cheap whereas engineering execution and tight feedback loops drive 90% of the actual value. | |
| History of Mechanistic Interpretability and the Superposition Hypothesis | 4 | 6 | 0 | 0 | The host and guest review the historical progression from Chris Olah's vision interpretability work to LLMs. Emmanuel explains the superposition hypothesis, contrasting the relatively small space of vision curves with the vast space of language concepts. | |
| Sparse Autoencoders, Dictionary Learning, and Golden Gate Claude | 5 | 6 | 0 | 0 | Emmanuel explains sparse autoencoders as unsupervised dictionary learning and recounts the origins of Golden Gate Claude. The host shares his own experiment building a Golden Gate Gemma by manipulating related concept activations. | |
| Feature Steering Trade-Offs, Model Safety, and Jailbreak Mechanisms | 5 | 6 | 1 | 3 | The host questions whether feature steering offers a free lunch for model quality. Emmanuel clarifies that interpretability is long-term research, illustrating its current utility through a jailbreak analysis where grammatical momentum overrides safety boundaries. | |
| AI Safety Imperatives and the Discovery of Induction Heads | 5 | 6 | 0 | 0 | The discussion covers safety imperatives and the mechanism of induction heads. The host connects induction heads to practical edit modes and code generation interfaces. | |
| Circuit Tracing Discoveries in Multi-Step Reasoning and Diagnostics | 5 | 7 | 0 | 2 | Emmanuel walks through internal multi-step reasoning examples (Dallas/Texas/Austin and medical test selection) to counter the 'stochastic parrot' narrative. The host asks about trade-offs between model depth and inference latency. | |
| Cross-Lingual Feature Sharing and Multimodal Concept Representations | 6 | 6 | 0 | 0 | Emmanuel explains how larger models share abstract concepts across human languages and modalities. The host demonstrates domain knowledge by connecting these findings to the Sapir-Whorf linguistic hypothesis and massive-context language learning. | |
| Internal Backward Planning in Poetry and Feature Labeling Methodology | 4 | 7 | 0 | 1 | Emmanuel presents findings on backward planning in poetry generation, demonstrating that models select rhyming targets before composing intervening lines. He also details the methodology behind labeling and verifying features. | |
| Mechanics of Attribution Graphs and Open Challenges in Interpretability | 4 | 7 | 0 | 0 | Emmanuel explains the mathematics of attribution graphs using backpropagation dot products between feature activations. He outlines major unsolved frontiers in interpretability, including attention decomposition and global model architectures. | |
| Chain of Thought Faithfulness, Deceptive Reasoning, and Planning | 5 | 7 | 2 | 2 | Emmanuel demonstrates chain of thought unfaithfulness, showing how a model back-computes incorrect arithmetic to match a suggested hint. He offers a $100 bet that this motivated reasoning stems from pre-training rather than RL fine-tuning. | |
| Publication Ethics, Model Self-Awareness, and Safety Trade-Offs | 5 | 5 | 0 | 3 | The host raises safety and publication ethics concerns about whether publishing interpretability research aids model situational awareness or alignment faking. Emmanuel discusses Anthropic's benefit-risk evaluations and alignment experiments. | |
| Behind the Scenes of High-End Interactive Data Visualizations | 4 | 5 | 0 | 0 | The hosts inquire into the production effort behind Anthropic's interactive D3 data visualizations. Emmanuel describes the internal tooling built by specialist engineers that enables researchers to assemble interactive diagrams. | |
| Future Directions, Fundamental Blockers, and Final Reflections | 3 | 5 | 0 | 0 | Vibhu asks about core blockers over the next 5 to 10 years. Emmanuel reflects on the transition from toy models to production models and encourages researchers to enter the open interpretability ecosystem. |