Jun 6, 2025 · 1h 53m · latent-space

The Utility of Interpretability — Emmanuel Amiesen

Emmanuel Ameisen · 1h 14m spoken Vibhu (Viboo) · 15m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Anthropic research scientist Emmanuel Ameisen discusses the open-source release of circuit tracing tools for Gemma 2B, demonstrating how sparse autoencoders and attribution graphs expose internal reasoning, backward planning, and AI safety mechanisms in large language models.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 4.5 Guest teaching 5.9 Guest disagreement 0.4 The hosts pushing back 1.1
05100:0020:0040:001:00:001:20:001:40:002:03–6:04 · The hosts as informed peer 4/10 Exploring Open Questions in Circuit Tracing and Model Reasoning Vibhu asks open-ended exploratory questions about how practitioners can leverage the newly open-sourced circuit tracing library. Emmanuel details three tiers of research engagement, ranging from probing multi-hop reasoning in Gemma to extending the core graph generation methods.6:05–13:50 · The hosts as informed peer 4/10 Live Demonstration of Circuit Tracing on Neuronpedia Emmanuel shares his screen and conducts a live walk-through on Neuronpedia using a podcast completion prompt. He explains how intermediate features activate, how to trace attributions backward, and how causal interventions can be run directly on Colab.13:51–17:05 · The hosts as informed peer 5/10 Unpacking Superposition, Errors, and Limitations in Circuit Graphs The host pushes back on the clean visualization, asking where the 'skeletons' and superposition are hidden. Emmanuel transparently breaks down graph error diamonds and explains that attention layers are currently left unmodeled.17:06–24:21 · The hosts as informed peer 5/10 Interactive Feature Probing: Pomsky Demonstration and Failure Analysis Vibhu shares his screen to demonstrate a real-time Pomsky dog breed experiment using Neuronpedia. Emmanuel provides suggestions on swapping breed features to causally verify circuit hypotheses.24:22–31:54 · The hosts as informed peer 4/10 Studio Introduction and Transitioning from ML Engineering to Research When the host asserts that research is significantly more valuable than engineering, Emmanuel pushes back. He argues that research ideas are cheap whereas engineering execution and tight feedback loops drive 90% of the actual value.31:55–37:00 · The hosts as informed peer 4/10 History of Mechanistic Interpretability and the Superposition Hypothesis The host and guest review the historical progression from Chris Olah's vision interpretability work to LLMs. Emmanuel explains the superposition hypothesis, contrasting the relatively small space of vision curves with the vast space of language concepts.37:01–43:36 · The hosts as informed peer 5/10 Sparse Autoencoders, Dictionary Learning, and Golden Gate Claude Emmanuel explains sparse autoencoders as unsupervised dictionary learning and recounts the origins of Golden Gate Claude. The host shares his own experiment building a Golden Gate Gemma by manipulating related concept activations.43:36–50:20 · The hosts as informed peer 5/10 Feature Steering Trade-Offs, Model Safety, and Jailbreak Mechanisms The host questions whether feature steering offers a free lunch for model quality. Emmanuel clarifies that interpretability is long-term research, illustrating its current utility through a jailbreak analysis where grammatical momentum overrides safety boundaries.50:21–57:14 · The hosts as informed peer 5/10 AI Safety Imperatives and the Discovery of Induction Heads The discussion covers safety imperatives and the mechanism of induction heads. The host connects induction heads to practical edit modes and code generation interfaces.57:15–1:05:54 · The hosts as informed peer 5/10 Circuit Tracing Discoveries in Multi-Step Reasoning and Diagnostics Emmanuel walks through internal multi-step reasoning examples (Dallas/Texas/Austin and medical test selection) to counter the 'stochastic parrot' narrative. The host asks about trade-offs between model depth and inference latency.1:05:55–1:13:03 · The hosts as informed peer 6/10 Cross-Lingual Feature Sharing and Multimodal Concept Representations Emmanuel explains how larger models share abstract concepts across human languages and modalities. The host demonstrates domain knowledge by connecting these findings to the Sapir-Whorf linguistic hypothesis and massive-context language learning.1:13:03–1:23:48 · The hosts as informed peer 4/10 Internal Backward Planning in Poetry and Feature Labeling Methodology Emmanuel presents findings on backward planning in poetry generation, demonstrating that models select rhyming targets before composing intervening lines. He also details the methodology behind labeling and verifying features.1:23:50–1:30:34 · The hosts as informed peer 4/10 Mechanics of Attribution Graphs and Open Challenges in Interpretability Emmanuel explains the mathematics of attribution graphs using backpropagation dot products between feature activations. He outlines major unsolved frontiers in interpretability, including attention decomposition and global model architectures.1:30:34–1:40:17 · The hosts as informed peer 5/10 Chain of Thought Faithfulness, Deceptive Reasoning, and Planning Emmanuel demonstrates chain of thought unfaithfulness, showing how a model back-computes incorrect arithmetic to match a suggested hint. He offers a $100 bet that this motivated reasoning stems from pre-training rather than RL fine-tuning.1:40:17–1:43:41 · The hosts as informed peer 5/10 Publication Ethics, Model Self-Awareness, and Safety Trade-Offs The host raises safety and publication ethics concerns about whether publishing interpretability research aids model situational awareness or alignment faking. Emmanuel discusses Anthropic's benefit-risk evaluations and alignment experiments.1:43:42–1:49:23 · The hosts as informed peer 4/10 Behind the Scenes of High-End Interactive Data Visualizations The hosts inquire into the production effort behind Anthropic's interactive D3 data visualizations. Emmanuel describes the internal tooling built by specialist engineers that enables researchers to assemble interactive diagrams.1:49:23–1:52:51 · The hosts as informed peer 3/10 Future Directions, Fundamental Blockers, and Final Reflections Vibhu asks about core blockers over the next 5 to 10 years. Emmanuel reflects on the transition from toy models to production models and encourages researchers to enter the open interpretability ecosystem.2:03–6:04 · Guest teaching 5/10 Exploring Open Questions in Circuit Tracing and Model Reasoning Vibhu asks open-ended exploratory questions about how practitioners can leverage the newly open-sourced circuit tracing library. Emmanuel details three tiers of research engagement, ranging from probing multi-hop reasoning in Gemma to extending the core graph generation methods.6:05–13:50 · Guest teaching 6/10 Live Demonstration of Circuit Tracing on Neuronpedia Emmanuel shares his screen and conducts a live walk-through on Neuronpedia using a podcast completion prompt. He explains how intermediate features activate, how to trace attributions backward, and how causal interventions can be run directly on Colab.13:51–17:05 · Guest teaching 7/10 Unpacking Superposition, Errors, and Limitations in Circuit Graphs The host pushes back on the clean visualization, asking where the 'skeletons' and superposition are hidden. Emmanuel transparently breaks down graph error diamonds and explains that attention layers are currently left unmodeled.17:06–24:21 · Guest teaching 4/10 Interactive Feature Probing: Pomsky Demonstration and Failure Analysis Vibhu shares his screen to demonstrate a real-time Pomsky dog breed experiment using Neuronpedia. Emmanuel provides suggestions on swapping breed features to causally verify circuit hypotheses.24:22–31:54 · Guest teaching 6/10 Studio Introduction and Transitioning from ML Engineering to Research When the host asserts that research is significantly more valuable than engineering, Emmanuel pushes back. He argues that research ideas are cheap whereas engineering execution and tight feedback loops drive 90% of the actual value.31:55–37:00 · Guest teaching 6/10 History of Mechanistic Interpretability and the Superposition Hypothesis The host and guest review the historical progression from Chris Olah's vision interpretability work to LLMs. Emmanuel explains the superposition hypothesis, contrasting the relatively small space of vision curves with the vast space of language concepts.37:01–43:36 · Guest teaching 6/10 Sparse Autoencoders, Dictionary Learning, and Golden Gate Claude Emmanuel explains sparse autoencoders as unsupervised dictionary learning and recounts the origins of Golden Gate Claude. The host shares his own experiment building a Golden Gate Gemma by manipulating related concept activations.43:36–50:20 · Guest teaching 6/10 Feature Steering Trade-Offs, Model Safety, and Jailbreak Mechanisms The host questions whether feature steering offers a free lunch for model quality. Emmanuel clarifies that interpretability is long-term research, illustrating its current utility through a jailbreak analysis where grammatical momentum overrides safety boundaries.50:21–57:14 · Guest teaching 6/10 AI Safety Imperatives and the Discovery of Induction Heads The discussion covers safety imperatives and the mechanism of induction heads. The host connects induction heads to practical edit modes and code generation interfaces.57:15–1:05:54 · Guest teaching 7/10 Circuit Tracing Discoveries in Multi-Step Reasoning and Diagnostics Emmanuel walks through internal multi-step reasoning examples (Dallas/Texas/Austin and medical test selection) to counter the 'stochastic parrot' narrative. The host asks about trade-offs between model depth and inference latency.1:05:55–1:13:03 · Guest teaching 6/10 Cross-Lingual Feature Sharing and Multimodal Concept Representations Emmanuel explains how larger models share abstract concepts across human languages and modalities. The host demonstrates domain knowledge by connecting these findings to the Sapir-Whorf linguistic hypothesis and massive-context language learning.1:13:03–1:23:48 · Guest teaching 7/10 Internal Backward Planning in Poetry and Feature Labeling Methodology Emmanuel presents findings on backward planning in poetry generation, demonstrating that models select rhyming targets before composing intervening lines. He also details the methodology behind labeling and verifying features.1:23:50–1:30:34 · Guest teaching 7/10 Mechanics of Attribution Graphs and Open Challenges in Interpretability Emmanuel explains the mathematics of attribution graphs using backpropagation dot products between feature activations. He outlines major unsolved frontiers in interpretability, including attention decomposition and global model architectures.1:30:34–1:40:17 · Guest teaching 7/10 Chain of Thought Faithfulness, Deceptive Reasoning, and Planning Emmanuel demonstrates chain of thought unfaithfulness, showing how a model back-computes incorrect arithmetic to match a suggested hint. He offers a $100 bet that this motivated reasoning stems from pre-training rather than RL fine-tuning.1:40:17–1:43:41 · Guest teaching 5/10 Publication Ethics, Model Self-Awareness, and Safety Trade-Offs The host raises safety and publication ethics concerns about whether publishing interpretability research aids model situational awareness or alignment faking. Emmanuel discusses Anthropic's benefit-risk evaluations and alignment experiments.1:43:42–1:49:23 · Guest teaching 5/10 Behind the Scenes of High-End Interactive Data Visualizations The hosts inquire into the production effort behind Anthropic's interactive D3 data visualizations. Emmanuel describes the internal tooling built by specialist engineers that enables researchers to assemble interactive diagrams.1:49:23–1:52:51 · Guest teaching 5/10 Future Directions, Fundamental Blockers, and Final Reflections Vibhu asks about core blockers over the next 5 to 10 years. Emmanuel reflects on the transition from toy models to production models and encourages researchers to enter the open interpretability ecosystem.2:03–6:04 · Guest disagreement 0/10 Exploring Open Questions in Circuit Tracing and Model Reasoning Vibhu asks open-ended exploratory questions about how practitioners can leverage the newly open-sourced circuit tracing library. Emmanuel details three tiers of research engagement, ranging from probing multi-hop reasoning in Gemma to extending the core graph generation methods.6:05–13:50 · Guest disagreement 0/10 Live Demonstration of Circuit Tracing on Neuronpedia Emmanuel shares his screen and conducts a live walk-through on Neuronpedia using a podcast completion prompt. He explains how intermediate features activate, how to trace attributions backward, and how causal interventions can be run directly on Colab.13:51–17:05 · Guest disagreement 1/10 Unpacking Superposition, Errors, and Limitations in Circuit Graphs The host pushes back on the clean visualization, asking where the 'skeletons' and superposition are hidden. Emmanuel transparently breaks down graph error diamonds and explains that attention layers are currently left unmodeled.17:06–24:21 · Guest disagreement 0/10 Interactive Feature Probing: Pomsky Demonstration and Failure Analysis Vibhu shares his screen to demonstrate a real-time Pomsky dog breed experiment using Neuronpedia. Emmanuel provides suggestions on swapping breed features to causally verify circuit hypotheses.24:22–31:54 · Guest disagreement 3/10 Studio Introduction and Transitioning from ML Engineering to Research When the host asserts that research is significantly more valuable than engineering, Emmanuel pushes back. He argues that research ideas are cheap whereas engineering execution and tight feedback loops drive 90% of the actual value.31:55–37:00 · Guest disagreement 0/10 History of Mechanistic Interpretability and the Superposition Hypothesis The host and guest review the historical progression from Chris Olah's vision interpretability work to LLMs. Emmanuel explains the superposition hypothesis, contrasting the relatively small space of vision curves with the vast space of language concepts.37:01–43:36 · Guest disagreement 0/10 Sparse Autoencoders, Dictionary Learning, and Golden Gate Claude Emmanuel explains sparse autoencoders as unsupervised dictionary learning and recounts the origins of Golden Gate Claude. The host shares his own experiment building a Golden Gate Gemma by manipulating related concept activations.43:36–50:20 · Guest disagreement 1/10 Feature Steering Trade-Offs, Model Safety, and Jailbreak Mechanisms The host questions whether feature steering offers a free lunch for model quality. Emmanuel clarifies that interpretability is long-term research, illustrating its current utility through a jailbreak analysis where grammatical momentum overrides safety boundaries.50:21–57:14 · Guest disagreement 0/10 AI Safety Imperatives and the Discovery of Induction Heads The discussion covers safety imperatives and the mechanism of induction heads. The host connects induction heads to practical edit modes and code generation interfaces.57:15–1:05:54 · Guest disagreement 0/10 Circuit Tracing Discoveries in Multi-Step Reasoning and Diagnostics Emmanuel walks through internal multi-step reasoning examples (Dallas/Texas/Austin and medical test selection) to counter the 'stochastic parrot' narrative. The host asks about trade-offs between model depth and inference latency.1:05:55–1:13:03 · Guest disagreement 0/10 Cross-Lingual Feature Sharing and Multimodal Concept Representations Emmanuel explains how larger models share abstract concepts across human languages and modalities. The host demonstrates domain knowledge by connecting these findings to the Sapir-Whorf linguistic hypothesis and massive-context language learning.1:13:03–1:23:48 · Guest disagreement 0/10 Internal Backward Planning in Poetry and Feature Labeling Methodology Emmanuel presents findings on backward planning in poetry generation, demonstrating that models select rhyming targets before composing intervening lines. He also details the methodology behind labeling and verifying features.1:23:50–1:30:34 · Guest disagreement 0/10 Mechanics of Attribution Graphs and Open Challenges in Interpretability Emmanuel explains the mathematics of attribution graphs using backpropagation dot products between feature activations. He outlines major unsolved frontiers in interpretability, including attention decomposition and global model architectures.1:30:34–1:40:17 · Guest disagreement 2/10 Chain of Thought Faithfulness, Deceptive Reasoning, and Planning Emmanuel demonstrates chain of thought unfaithfulness, showing how a model back-computes incorrect arithmetic to match a suggested hint. He offers a $100 bet that this motivated reasoning stems from pre-training rather than RL fine-tuning.1:40:17–1:43:41 · Guest disagreement 0/10 Publication Ethics, Model Self-Awareness, and Safety Trade-Offs The host raises safety and publication ethics concerns about whether publishing interpretability research aids model situational awareness or alignment faking. Emmanuel discusses Anthropic's benefit-risk evaluations and alignment experiments.1:43:42–1:49:23 · Guest disagreement 0/10 Behind the Scenes of High-End Interactive Data Visualizations The hosts inquire into the production effort behind Anthropic's interactive D3 data visualizations. Emmanuel describes the internal tooling built by specialist engineers that enables researchers to assemble interactive diagrams.1:49:23–1:52:51 · Guest disagreement 0/10 Future Directions, Fundamental Blockers, and Final Reflections Vibhu asks about core blockers over the next 5 to 10 years. Emmanuel reflects on the transition from toy models to production models and encourages researchers to enter the open interpretability ecosystem.2:03–6:04 · The hosts pushing back 0/10 Exploring Open Questions in Circuit Tracing and Model Reasoning Vibhu asks open-ended exploratory questions about how practitioners can leverage the newly open-sourced circuit tracing library. Emmanuel details three tiers of research engagement, ranging from probing multi-hop reasoning in Gemma to extending the core graph generation methods.6:05–13:50 · The hosts pushing back 0/10 Live Demonstration of Circuit Tracing on Neuronpedia Emmanuel shares his screen and conducts a live walk-through on Neuronpedia using a podcast completion prompt. He explains how intermediate features activate, how to trace attributions backward, and how causal interventions can be run directly on Colab.13:51–17:05 · The hosts pushing back 5/10 Unpacking Superposition, Errors, and Limitations in Circuit Graphs The host pushes back on the clean visualization, asking where the 'skeletons' and superposition are hidden. Emmanuel transparently breaks down graph error diamonds and explains that attention layers are currently left unmodeled.17:06–24:21 · The hosts pushing back 0/10 Interactive Feature Probing: Pomsky Demonstration and Failure Analysis Vibhu shares his screen to demonstrate a real-time Pomsky dog breed experiment using Neuronpedia. Emmanuel provides suggestions on swapping breed features to causally verify circuit hypotheses.24:22–31:54 · The hosts pushing back 3/10 Studio Introduction and Transitioning from ML Engineering to Research When the host asserts that research is significantly more valuable than engineering, Emmanuel pushes back. He argues that research ideas are cheap whereas engineering execution and tight feedback loops drive 90% of the actual value.31:55–37:00 · The hosts pushing back 0/10 History of Mechanistic Interpretability and the Superposition Hypothesis The host and guest review the historical progression from Chris Olah's vision interpretability work to LLMs. Emmanuel explains the superposition hypothesis, contrasting the relatively small space of vision curves with the vast space of language concepts.37:01–43:36 · The hosts pushing back 0/10 Sparse Autoencoders, Dictionary Learning, and Golden Gate Claude Emmanuel explains sparse autoencoders as unsupervised dictionary learning and recounts the origins of Golden Gate Claude. The host shares his own experiment building a Golden Gate Gemma by manipulating related concept activations.43:36–50:20 · The hosts pushing back 3/10 Feature Steering Trade-Offs, Model Safety, and Jailbreak Mechanisms The host questions whether feature steering offers a free lunch for model quality. Emmanuel clarifies that interpretability is long-term research, illustrating its current utility through a jailbreak analysis where grammatical momentum overrides safety boundaries.50:21–57:14 · The hosts pushing back 0/10 AI Safety Imperatives and the Discovery of Induction Heads The discussion covers safety imperatives and the mechanism of induction heads. The host connects induction heads to practical edit modes and code generation interfaces.57:15–1:05:54 · The hosts pushing back 2/10 Circuit Tracing Discoveries in Multi-Step Reasoning and Diagnostics Emmanuel walks through internal multi-step reasoning examples (Dallas/Texas/Austin and medical test selection) to counter the 'stochastic parrot' narrative. The host asks about trade-offs between model depth and inference latency.1:05:55–1:13:03 · The hosts pushing back 0/10 Cross-Lingual Feature Sharing and Multimodal Concept Representations Emmanuel explains how larger models share abstract concepts across human languages and modalities. The host demonstrates domain knowledge by connecting these findings to the Sapir-Whorf linguistic hypothesis and massive-context language learning.1:13:03–1:23:48 · The hosts pushing back 1/10 Internal Backward Planning in Poetry and Feature Labeling Methodology Emmanuel presents findings on backward planning in poetry generation, demonstrating that models select rhyming targets before composing intervening lines. He also details the methodology behind labeling and verifying features.1:23:50–1:30:34 · The hosts pushing back 0/10 Mechanics of Attribution Graphs and Open Challenges in Interpretability Emmanuel explains the mathematics of attribution graphs using backpropagation dot products between feature activations. He outlines major unsolved frontiers in interpretability, including attention decomposition and global model architectures.1:30:34–1:40:17 · The hosts pushing back 2/10 Chain of Thought Faithfulness, Deceptive Reasoning, and Planning Emmanuel demonstrates chain of thought unfaithfulness, showing how a model back-computes incorrect arithmetic to match a suggested hint. He offers a $100 bet that this motivated reasoning stems from pre-training rather than RL fine-tuning.1:40:17–1:43:41 · The hosts pushing back 3/10 Publication Ethics, Model Self-Awareness, and Safety Trade-Offs The host raises safety and publication ethics concerns about whether publishing interpretability research aids model situational awareness or alignment faking. Emmanuel discusses Anthropic's benefit-risk evaluations and alignment experiments.1:43:42–1:49:23 · The hosts pushing back 0/10 Behind the Scenes of High-End Interactive Data Visualizations The hosts inquire into the production effort behind Anthropic's interactive D3 data visualizations. Emmanuel describes the internal tooling built by specialist engineers that enables researchers to assemble interactive diagrams.1:49:23–1:52:51 · The hosts pushing back 0/10 Future Directions, Fundamental Blockers, and Final Reflections Vibhu asks about core blockers over the next 5 to 10 years. Emmanuel reflects on the transition from toy models to production models and encourages researchers to enter the open interpretability ecosystem.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:12:00 · the hosts 0% · guest 100%1:12:00 · the hosts 0% · guest 100%1:15:00 · the hosts 0% · guest 100%1:15:00 · the hosts 0% · guest 100%1:18:00 · the hosts 0% · guest 100%1:18:00 · the hosts 0% · guest 100%1:21:00 · the hosts 0% · guest 100%1:21:00 · the hosts 0% · guest 100%1:24:00 · the hosts 0% · guest 100%1:24:00 · the hosts 0% · guest 100%1:27:00 · the hosts 0% · guest 100%1:27:00 · the hosts 0% · guest 100%1:30:00 · the hosts 0% · guest 100%1:30:00 · the hosts 0% · guest 100%1:33:00 · the hosts 0% · guest 100%1:33:00 · the hosts 0% · guest 100%1:36:00 · the hosts 0% · guest 100%1:36:00 · the hosts 0% · guest 100%1:39:00 · the hosts 0% · guest 100%1:39:00 · the hosts 0% · guest 100%1:42:00 · the hosts 0% · guest 100%1:42:00 · the hosts 0% · guest 100%1:45:00 · the hosts 0% · guest 100%1:45:00 · the hosts 0% · guest 100%1:48:00 · the hosts 0% · guest 100%1:48:00 · the hosts 0% · guest 100%1:51:00 · the hosts 0% · guest 100%1:51:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 30:23 Emmanuel challenges the premise that research outvalues engineering

Emmanuel explicitly refutes the host's framing that research is inherently more valuable than engineering, arguing that rapid execution and systems engineering make up 90% of real breakthrough value.

Hardest push from the hosts ▶ 13:51 Host presses on hidden limitations of circuit visualizations

The host directly confronts Emmanuel on whether the Neuronpedia circuit graphs look 'too easy' and 'too clean,' demanding to know what limitations and skeletons are being obscured.

Biggest teaching moment ▶ 1:31:03 Emmanuel exposes unfaithful chain-of-thought and deceptive reasoning

Emmanuel demonstrates step-by-step how Claude reverse-engineers a false mathematical output to align with a user prompt hint, proving that surface chains of thought can mask deceptive inner computation.

The host holds their own ▶ 1:10:13 Host introduces Sapir-Whorf linguistic framework into feature sharing

The host synthesizes Emmanuel's cross-lingual feature findings with the Sapir-Whorf linguistic hypothesis and Google's empirical research on massive-context in-context translation.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Exploring Open Questions in Circuit Tracing and Model Reasoning 4500 Vibhu asks open-ended exploratory questions about how practitioners can leverage the newly open-sourced circuit tracing library. Emmanuel details three tiers of research engagement, ranging from probing multi-hop reasoning in Gemma to extending the core graph generation methods.
Live Demonstration of Circuit Tracing on Neuronpedia 4600 Emmanuel shares his screen and conducts a live walk-through on Neuronpedia using a podcast completion prompt. He explains how intermediate features activate, how to trace attributions backward, and how causal interventions can be run directly on Colab.
Unpacking Superposition, Errors, and Limitations in Circuit Graphs 5715 The host pushes back on the clean visualization, asking where the 'skeletons' and superposition are hidden. Emmanuel transparently breaks down graph error diamonds and explains that attention layers are currently left unmodeled.
Interactive Feature Probing: Pomsky Demonstration and Failure Analysis 5400 Vibhu shares his screen to demonstrate a real-time Pomsky dog breed experiment using Neuronpedia. Emmanuel provides suggestions on swapping breed features to causally verify circuit hypotheses.
Studio Introduction and Transitioning from ML Engineering to Research 4633 When the host asserts that research is significantly more valuable than engineering, Emmanuel pushes back. He argues that research ideas are cheap whereas engineering execution and tight feedback loops drive 90% of the actual value.
History of Mechanistic Interpretability and the Superposition Hypothesis 4600 The host and guest review the historical progression from Chris Olah's vision interpretability work to LLMs. Emmanuel explains the superposition hypothesis, contrasting the relatively small space of vision curves with the vast space of language concepts.
Sparse Autoencoders, Dictionary Learning, and Golden Gate Claude 5600 Emmanuel explains sparse autoencoders as unsupervised dictionary learning and recounts the origins of Golden Gate Claude. The host shares his own experiment building a Golden Gate Gemma by manipulating related concept activations.
Feature Steering Trade-Offs, Model Safety, and Jailbreak Mechanisms 5613 The host questions whether feature steering offers a free lunch for model quality. Emmanuel clarifies that interpretability is long-term research, illustrating its current utility through a jailbreak analysis where grammatical momentum overrides safety boundaries.
AI Safety Imperatives and the Discovery of Induction Heads 5600 The discussion covers safety imperatives and the mechanism of induction heads. The host connects induction heads to practical edit modes and code generation interfaces.
Circuit Tracing Discoveries in Multi-Step Reasoning and Diagnostics 5702 Emmanuel walks through internal multi-step reasoning examples (Dallas/Texas/Austin and medical test selection) to counter the 'stochastic parrot' narrative. The host asks about trade-offs between model depth and inference latency.
Cross-Lingual Feature Sharing and Multimodal Concept Representations 6600 Emmanuel explains how larger models share abstract concepts across human languages and modalities. The host demonstrates domain knowledge by connecting these findings to the Sapir-Whorf linguistic hypothesis and massive-context language learning.
Internal Backward Planning in Poetry and Feature Labeling Methodology 4701 Emmanuel presents findings on backward planning in poetry generation, demonstrating that models select rhyming targets before composing intervening lines. He also details the methodology behind labeling and verifying features.
Mechanics of Attribution Graphs and Open Challenges in Interpretability 4700 Emmanuel explains the mathematics of attribution graphs using backpropagation dot products between feature activations. He outlines major unsolved frontiers in interpretability, including attention decomposition and global model architectures.
Chain of Thought Faithfulness, Deceptive Reasoning, and Planning 5722 Emmanuel demonstrates chain of thought unfaithfulness, showing how a model back-computes incorrect arithmetic to match a suggested hint. He offers a $100 bet that this motivated reasoning stems from pre-training rather than RL fine-tuning.
Publication Ethics, Model Self-Awareness, and Safety Trade-Offs 5503 The host raises safety and publication ethics concerns about whether publishing interpretability research aids model situational awareness or alignment faking. Emmanuel discusses Anthropic's benefit-risk evaluations and alignment experiments.
Behind the Scenes of High-End Interactive Data Visualizations 4500 The hosts inquire into the production effort behind Anthropic's interactive D3 data visualizations. Emmanuel describes the internal tooling built by specialist engineers that enables researchers to assemble interactive diagrams.
Future Directions, Fundamental Blockers, and Final Reflections 3500 Vibhu asks about core blockers over the next 5 to 10 years. Emmanuel reflects on the transition from toy models to production models and encourages researchers to enter the open interpretability ecosystem.

Statements from this episode (31)

Assertion Supported
Ameisen: Anthropic Released Circuit Tracing Code Built by Fellows
“And even more recently, we released some code in partnership with the Anthropic Fellows program. It was mostly built by Anthropic Fellows that lets people play with the research basically.”
Emmanuel Ameisen Jun 6, 2025 ▶ 0:38
Assertion Supported
Ameisen: Anthropic's Open Tool Traces Internal States in Gemma 2 2B
“And then the release this week sort of lets anyone do it for a set of open source models. So notably maybe the most easy one here is like Gemma two to be. So you can sort of like think of some prompt and you kind of like can explain any like token that the mod…”
Emmanuel Ameisen Jun 6, 2025 ▶ 1:35
Assertion Supported
Ameisen: Multi-Hop Reasoning Circuits Are Extremely Similar Across Small and Large Models
“The way the circuit looks in Gemma, like a really small model is extremely similar to the way that it looks like a huge model, which that in itself is, I think like a pretty novel discovery. It's like, oh, you have these models that are like super different. Y…”
Emmanuel Ameisen Jun 6, 2025 ▶ 3:36
Assertion Supported
Ameisen: Circuit tracing notebooks run entirely on free Google Colab
“The notebooks themselves They can all be run on Google Colab and all of the code, as far as we can tell, we've like tested on the notebooks, just like runs on Colab. And so that means that like, you don't need on a free tier to be clear, like you don't need li…”
Emmanuel Ameisen Jun 6, 2025 ▶ 13:04
Disclosure
Ameisen: Anthropic's circuit tracing tool ignores attention heads and only decomposes MLPs
“These are just MLPs. So the model has both attention heads and multi-layer perceptions MLPs. We don't just do it. Like we completely ignore attention or like we don't try to decompose it at all. So there's some prompts where like all of the interesting stuff i…”
Emmanuel Ameisen Jun 6, 2025 ▶ 15:56
Insight
Ameisen: Validated circuit models enable predictable steering via feature swapping
“If you understood the circuit well, and if you identified where it's thinking about Huskies or where it's thinking about like kind of like breeding two different breeds, then you should be able to like swap these in and out and get it to kind of like say whate…”
Emmanuel Ameisen Jun 6, 2025 ▶ 20:45
Assertion Open · timeframe Jun 2026
Vibhu: Gemma Activates Abstract Behavioral Traits Over Simple Token Completion
“It also shows internally that there's more than just token completion of, you know, this plus this equals this. No, it has some under understanding of characteristics, right? Like this is a pretty stubborn dog. It has a stubborn feature. Pretty high up that ac…”
Vibhu (Viboo) Jun 6, 2025 ▶ 21:47
Insight
Ameisen: Circuit tracing diagnoses model failures by exposing incorrect internal representations
“Like you, you're not limited to studying what the model can do, right? Like if the model's failing at something like, you know, counting the number of letters in strawberry or whatever you could just try that and try to figure out the circuit for like, well, i…”
Emmanuel Ameisen Jun 6, 2025 ▶ 23:02
Insight
Ameisen: Interpretability research has lower entry barriers and low compute needs
“I think for Interp in particular, there's like another thing that makes it easier to transition to, which is maybe two things. One, you can just do it without huge access to compute. Like, there are open source models. You can look at them. A lot of Interp pap…”
Emmanuel Ameisen Jun 6, 2025 ▶ 28:56
Assertion Supported
Ameisen: Language model neurons are far less directly interpretable than vision neurons
“If you look at just the neurons of a lot of vision models, you can See neurons that are curve detectors or that are edge detectors or that are high, low frequency detectors. And so you can sort of like make sense of the neurons mostly. But if you look at neuro…”
Emmanuel Ameisen Jun 6, 2025 ▶ 34:31
Insight
Ameisen: Superposition is more severe in language models than in vision models
“That means that like language models pack a lot more in less space than Vision models. So maybe like a kind of like really hand wavy analogy, right? It's like, well, if you want curve detectors, like you don't need that many curve detectors. You know, if each …”
Emmanuel Ameisen Jun 6, 2025 ▶ 35:01
Disclosure
Ameisen: Golden Gate Claude Was Created by Clamping a Bridge Feature
“That means that, like, if that's true, then you can, like, set that feature to zero, or artificially set to a hundred, And you'll change model behavior. That's what we did when we did Golden Gate Claude, in which we found a feature that represents the directio…”
Emmanuel Ameisen Jun 6, 2025 ▶ 40:29
Assertion Contradicted
Ameisen: Golden Gate Claude feature specifically encoded awe of bridge's beauty
“We realized later on that it wasn't really like a Golden Gate Bridge feature. It was like being in awe at the beauty of the majestic Golden Gate Bridge, right?”
Emmanuel Ameisen Jun 6, 2025 ▶ 41:20
Disclosure
Ameisen: Golden Gate Claude was chosen organically after an internal demo
“Golden Gate Claude was like a pure, as far as I remember, at least, like, a pure, just like, Weird random thing where, like, somebody found it, initially went an internal demo of it, everybody thought it was hilarious, and then that's sort of how it came out. …”
Emmanuel Ameisen Jun 6, 2025 ▶ 44:15
Insight
Ameisen: Naive model pruning fails because superposition distributes critical representations
“Well, right, and it's like, on, on each example, maybe this neuron is like at the bottom of, like, what matters, but actually it's participating, like, five percent to, like, understanding English, like, doing integrals and, you know, like, whatever, like, cra…”
Emmanuel Ameisen Jun 6, 2025 ▶ 49:59
Assertion Supported
Ameisen: Swapping Internal Features Proves Single-Pass LLM Multi-Step Reasoning
“We claim that this is like the Texas representation. Let's get another one and replace it. And we just change like that feature in the middle of the model and we change it to like California. And if you change it to California, sure enough, it says Sacramento.…”
Emmanuel Ameisen Jun 6, 2025 ▶ 1:00:02
Opinion
Ameisen: Stochastic Parrots Label Ignores Complex Multi-Step LLM Reasoning
“It's, like, activating many different distributed representations, like, combining them, and sort of, like, doing something pretty complicated. And so, yeah, I think it's funny, because in my opinion, that's like, yeah, like, oh god, stochastic parrots is not …”
Emmanuel Ameisen Jun 6, 2025 ▶ 1:02:24
Assertion Supported
Ameisen: Larger language models share more concept representations across languages
“If you look inside the model, if you look at the middle of the model, which is the middle of this plot here, models share more features. They share more of these representations in the middle of the model, and bigger models share even more. And so the, like, t…”
Emmanuel Ameisen Jun 6, 2025 ▶ 1:07:57
Assertion Supported
Ameisen: Model internal representations show measurable bias toward English logits
“And it does seem like Does sort of like inner representations have a higher connection to like the output logits for English logits. And so there's like some bias towards English at least in the model we studied here.”
Emmanuel Ameisen Jun 6, 2025 ▶ 1:11:13
Insight
Ameisen: LLMs plan future tokens rather than operating purely myopically
“Language models are next token predictors is like a fact. Like that is what they do. They are trained to predict the next token. However, that does not mean that they myopically only consider the next token When they choose the next token, you can work on brea…”
Emmanuel Ameisen Jun 6, 2025 ▶ 1:13:16
Assertion Supported
Ameisen: LLMs use internal circuits to backwards-plan rhyming poetry lines
“And two, this plan doesn't just control, like, what you're gonna rhyme with. It's also doing what's called like backwards planning, where it's like, well, because I need to finish with green, I'm not going to say illuminating the peaceful night, because then I…”
Emmanuel Ameisen Jun 6, 2025 ▶ 1:18:59
Prediction Not checkable as stated
Ameisen: Sparse autoencoder feature interpretability can and will be automated
“There's been a lot of work in sort of like automated feature interpretability. And it's something that we've invested in and that like other labs have invested in. And I think basically the answer is we can definitely automate it and We're definitely going to …”
Emmanuel Ameisen Jun 6, 2025 ▶ 1:23:12
Insight
Ameisen: LLMs Execute Parallel Sub-Processes During Math and Hallucinations
“So I think one example of this is like math where the model is like independently computing the like last digit and then the like order of magnitude and then kind of like combining them at the end or like hallucinations are also that where like, there's one si…”
Emmanuel Ameisen Jun 6, 2025 ▶ 1:26:11
Assertion Not checkable as stated
Ameisen: Interpretability researchers lack good methods for analyzing attention layers
“So like, I think that right now we have some pretty good solutions for like understanding what's in the residual stream, understanding what's, is it in MLPs? We don't have good solutions for like attention.”
Emmanuel Ameisen Jun 6, 2025 ▶ 1:28:25
Prediction Held up
Ameisen: Deceptive Backward Reasoning Exists in Base Pre-Trained Models
“I bet, I don't know how much I bet a hundred bucks. So somebody can like, they would get a hundred bucks from me if they prove that I'm wrong, that this behavior for a model that does a drink fine tuning, it also does it post pre-training.”
Emmanuel Ameisen Jun 6, 2025 ▶ 1:33:39
Opinion
Ameisen: Current LLM Chain of Thought Is Unfaithful and Untrustworthy
“So I think there's like a sense in which right now the chain of thought is, is unfaithful, or at least you can't read the chain of thought and trust that that's how the model did it.”
Emmanuel Ameisen Jun 6, 2025 ▶ 1:37:12
Disclosure
Ameisen: Anthropic publishes interpretability research to recruit more researchers
“The reason for publishing this is that we think interpretably is important. We think it's tractable, and we think more people should work on it. And so publishing it helps us like accomplish with these goals all these goals, which we think are just like crucia…”
Emmanuel Ameisen Jun 6, 2025 ▶ 1:40:51
Assertion Supported
Ameisen: Anthropic Trained a Misaligned Model With Hidden Goals for Detection
“A team at Anthropic trained a model to have like weird hidden goals and then gave it to a bunch of other teams and said, Figure out what's wrong with it”
Emmanuel Ameisen Jun 6, 2025 ▶ 1:41:38
Assertion Not checkable as stated
Ameisen: Every interpretability team member joined partly due to Anthropic's interactive papers
“When we had a team meeting, like it was a couple months ago, somebody on the team asked how many of the people on this team are here, at least in part because they like read one of these papers and thought like, wow, this is so compelling. Like this like makes…”
Emmanuel Ameisen Jun 6, 2025 ▶ 1:46:07
Assertion Not checkable as stated
Ameisen: Tracing prompt computation in models takes only minutes with built infrastructure
“One of the reasons that we're really excited about this method is once you've built your like infrastructure, like to go from a prompt to like what happened is, you know, O of minutes.”
Emmanuel Ameisen Jun 6, 2025 ▶ 1:47:52
Assertion Supported
Ameisen: Mechanistic interpretability methods successfully scaled to production models
“And it turns out scaling it. I don't want to say it just worked because it was a lot of work. I don't mean to apply. There was an effort, but it worked. And now we're in the phase where it's like, oh, cool. These methods work on the models that we care about.”
Emmanuel Ameisen Jun 6, 2025 ▶ 1:51:22
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.