Feb 5, 2026 · 1h 8m · latent-space

Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell

Mark Bissell · 26m spoken Myra Deng · 16m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Goodfire AI's Mark Bissell and Myra Deng join the Latent Space Podcast to discuss how mechanistic interpretability is evolving from passive model inspection into an active engineering discipline. They demonstrate real-time activation steering, enterprise safety guardrails, and applications spanning life sciences and foundation model architecture design.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 6.0 Guest teaching 5.3 Guest disagreement 0.7 The hosts pushing back 1.5
05100:0015:0030:0045:001:00:002:47–5:40 · The hosts as informed peer 4/10 Engineering Backgrounds and Early Team Roles at Goodfire The host and guests share background stories collegially, including overlapping history at Two Sigma and Palantir. The dynamic is conversational and welcoming.5:40–8:08 · The hosts as informed peer 7/10 Unpacking Mechanistic Interpretability Beyond Black-Box Methods Co-host Vibu gives a detailed overview of mechanistic interpretability techniques like SAEs and probing. Mark expands the scope from post-hoc poking to intentional architecture design during training.8:09–12:09 · The hosts as informed peer 7/10 Post-Training Interventions, Surgical Edits, and Grokking Dynamics The host engages deeply on grokking and double descent, suggesting MechInterp reframes double descent as translating representation space. Mark builds on this by explaining surgical edits and generalizing solutions.12:10–16:58 · The hosts as informed peer 6/10 Investigating Subliminal Learning and Shared Latent Geometry When discussing subliminal learning, the host posits a Platonic representation hypothesis. Mark counters that the phenomenon is primarily a path-dependent statistical artifact stemming from shared initial random seeds.16:58–19:04 · The hosts as informed peer 5/10 Evaluating Linear Probes Versus Sparse Autoencoders in Practice The host presses on the limitations of Sparse Autoencoders. Myra explains their counterintuitive finding where raw activation probes consistently outperform SAE probes except on noisy datasets.19:04–21:56 · The hosts as informed peer 6/10 Enterprise Deployment Case Study: Real-Time Guardrails at Rakuten Vibu notes the runtime latency advantage of internal probes over LLM-as-a-judge guardrails. Mark outlines Rakuten's real-world deployment challenges including synthetic-to-real transfer and token-level PII scrubbing.21:56–25:32 · The hosts as informed peer 4/10 Live Demonstration: Real-Time Activation Steering on Kimi K2 Mark showcases a live CLI demonstration steering activations on the 1T parameter Kimi K2 model. The host follows the live output transition into Gen Z slang.25:32–29:35 · The hosts as informed peer 6/10 Feature Extraction Pipelines, Labeling, and Hallucination Detection The host calls out the fallacy of setting temperature to zero to stop hallucinations. Mark and Myra explain why supervised probes are superior to unsupervised SAEs for hallucination detection due to feature splitting.29:35–33:42 · The hosts as informed peer 7/10 Mathematical Equivalence of Activation Steering and In-Context Learning Mark explains team research proving mathematical equivalence between activation steering and in-context learning. The host synthesizes this with KV cache state modifications during inference.33:42–37:21 · The hosts as informed peer 7/10 Moving Beyond Brute-Force RL Toward Intentional Model Architecture The host compares low-rank adapters to steering. Mark uses the metaphor of modifying the physical pipes versus modifying the water flowing through them to explain intentional architectural design over brute-force RL.37:21–42:25 · The hosts as informed peer 5/10 Accessibility, Community Infrastructure, and Educational Pathways in MechInterp The hosts and guests discuss the low compute barrier to entry in MechInterp research, educational pathways like MATS, and the launch of dedicated industry tracks at AI Engineer conferences.42:25–46:34 · The hosts as informed peer 6/10 Code Representations, Steering Limits, and Scaling Frontiers The host notes public skepticism that steering can move beyond stylistic quirks like Gen Z tone into deep reasoning. Myra acknowledges these limits and highlights the need for new learning algorithms beyond simple scale.46:35–52:20 · The hosts as informed peer 6/10 Interpretability for Life Sciences: Evo 2 and Biomarker Discovery Mark outlines bidirectional interpretability in life sciences models, explaining how Goodfire partnered with Mayo Clinic and Prima Menta to uncover novel biomarkers for Alzheimer's disease.52:20–57:10 · The hosts as informed peer 6/10 Applying Interpretability to Vision, Diffusion Models, and World Models Myra explains how pixel space offers immediate visual feedback loops for interpretability. Mark discusses how astrophysics models learn alien heuristics that functionally work without grokking standard physical laws.57:10–1:00:53 · The hosts as informed peer 7/10 Science Fiction Inspirations: Ted Chiang, Model Sandbagging, and Neurobiology The host connects Ted Chiang's story 'Understand' to LLM model sandbagging and superintelligence test-evasion. Mark references Chiang's 'Exhalation' as an allegory for self-interpretability.1:00:53–1:06:14 · The hosts as informed peer 7/10 Pragmatic AI Safety, Superalignment, and Scalable Oversight The host questions whether weak-to-strong generalization will scale once models surpass human ability. Myra and Mark emphasize pragmatic, grounded alignment through scalable oversight rather than doomerism.2:47–5:40 · Guest teaching 1/10 Engineering Backgrounds and Early Team Roles at Goodfire The host and guests share background stories collegially, including overlapping history at Two Sigma and Palantir. The dynamic is conversational and welcoming.5:40–8:08 · Guest teaching 4/10 Unpacking Mechanistic Interpretability Beyond Black-Box Methods Co-host Vibu gives a detailed overview of mechanistic interpretability techniques like SAEs and probing. Mark expands the scope from post-hoc poking to intentional architecture design during training.8:09–12:09 · Guest teaching 5/10 Post-Training Interventions, Surgical Edits, and Grokking Dynamics The host engages deeply on grokking and double descent, suggesting MechInterp reframes double descent as translating representation space. Mark builds on this by explaining surgical edits and generalizing solutions.12:10–16:58 · Guest teaching 6/10 Investigating Subliminal Learning and Shared Latent Geometry When discussing subliminal learning, the host posits a Platonic representation hypothesis. Mark counters that the phenomenon is primarily a path-dependent statistical artifact stemming from shared initial random seeds.16:58–19:04 · Guest teaching 6/10 Evaluating Linear Probes Versus Sparse Autoencoders in Practice The host presses on the limitations of Sparse Autoencoders. Myra explains their counterintuitive finding where raw activation probes consistently outperform SAE probes except on noisy datasets.19:04–21:56 · Guest teaching 5/10 Enterprise Deployment Case Study: Real-Time Guardrails at Rakuten Vibu notes the runtime latency advantage of internal probes over LLM-as-a-judge guardrails. Mark outlines Rakuten's real-world deployment challenges including synthetic-to-real transfer and token-level PII scrubbing.21:56–25:32 · Guest teaching 5/10 Live Demonstration: Real-Time Activation Steering on Kimi K2 Mark showcases a live CLI demonstration steering activations on the 1T parameter Kimi K2 model. The host follows the live output transition into Gen Z slang.25:32–29:35 · Guest teaching 6/10 Feature Extraction Pipelines, Labeling, and Hallucination Detection The host calls out the fallacy of setting temperature to zero to stop hallucinations. Mark and Myra explain why supervised probes are superior to unsupervised SAEs for hallucination detection due to feature splitting.29:35–33:42 · Guest teaching 7/10 Mathematical Equivalence of Activation Steering and In-Context Learning Mark explains team research proving mathematical equivalence between activation steering and in-context learning. The host synthesizes this with KV cache state modifications during inference.33:42–37:21 · Guest teaching 6/10 Moving Beyond Brute-Force RL Toward Intentional Model Architecture The host compares low-rank adapters to steering. Mark uses the metaphor of modifying the physical pipes versus modifying the water flowing through them to explain intentional architectural design over brute-force RL.37:21–42:25 · Guest teaching 4/10 Accessibility, Community Infrastructure, and Educational Pathways in MechInterp The hosts and guests discuss the low compute barrier to entry in MechInterp research, educational pathways like MATS, and the launch of dedicated industry tracks at AI Engineer conferences.42:25–46:34 · Guest teaching 6/10 Code Representations, Steering Limits, and Scaling Frontiers The host notes public skepticism that steering can move beyond stylistic quirks like Gen Z tone into deep reasoning. Myra acknowledges these limits and highlights the need for new learning algorithms beyond simple scale.46:35–52:20 · Guest teaching 7/10 Interpretability for Life Sciences: Evo 2 and Biomarker Discovery Mark outlines bidirectional interpretability in life sciences models, explaining how Goodfire partnered with Mayo Clinic and Prima Menta to uncover novel biomarkers for Alzheimer's disease.52:20–57:10 · Guest teaching 6/10 Applying Interpretability to Vision, Diffusion Models, and World Models Myra explains how pixel space offers immediate visual feedback loops for interpretability. Mark discusses how astrophysics models learn alien heuristics that functionally work without grokking standard physical laws.57:10–1:00:53 · Guest teaching 5/10 Science Fiction Inspirations: Ted Chiang, Model Sandbagging, and Neurobiology The host connects Ted Chiang's story 'Understand' to LLM model sandbagging and superintelligence test-evasion. Mark references Chiang's 'Exhalation' as an allegory for self-interpretability.1:00:53–1:06:14 · Guest teaching 6/10 Pragmatic AI Safety, Superalignment, and Scalable Oversight The host questions whether weak-to-strong generalization will scale once models surpass human ability. Myra and Mark emphasize pragmatic, grounded alignment through scalable oversight rather than doomerism.2:47–5:40 · Guest disagreement 0/10 Engineering Backgrounds and Early Team Roles at Goodfire The host and guests share background stories collegially, including overlapping history at Two Sigma and Palantir. The dynamic is conversational and welcoming.5:40–8:08 · Guest disagreement 1/10 Unpacking Mechanistic Interpretability Beyond Black-Box Methods Co-host Vibu gives a detailed overview of mechanistic interpretability techniques like SAEs and probing. Mark expands the scope from post-hoc poking to intentional architecture design during training.8:09–12:09 · Guest disagreement 1/10 Post-Training Interventions, Surgical Edits, and Grokking Dynamics The host engages deeply on grokking and double descent, suggesting MechInterp reframes double descent as translating representation space. Mark builds on this by explaining surgical edits and generalizing solutions.12:10–16:58 · Guest disagreement 2/10 Investigating Subliminal Learning and Shared Latent Geometry When discussing subliminal learning, the host posits a Platonic representation hypothesis. Mark counters that the phenomenon is primarily a path-dependent statistical artifact stemming from shared initial random seeds.16:58–19:04 · Guest disagreement 1/10 Evaluating Linear Probes Versus Sparse Autoencoders in Practice The host presses on the limitations of Sparse Autoencoders. Myra explains their counterintuitive finding where raw activation probes consistently outperform SAE probes except on noisy datasets.19:04–21:56 · Guest disagreement 0/10 Enterprise Deployment Case Study: Real-Time Guardrails at Rakuten Vibu notes the runtime latency advantage of internal probes over LLM-as-a-judge guardrails. Mark outlines Rakuten's real-world deployment challenges including synthetic-to-real transfer and token-level PII scrubbing.21:56–25:32 · Guest disagreement 0/10 Live Demonstration: Real-Time Activation Steering on Kimi K2 Mark showcases a live CLI demonstration steering activations on the 1T parameter Kimi K2 model. The host follows the live output transition into Gen Z slang.25:32–29:35 · Guest disagreement 1/10 Feature Extraction Pipelines, Labeling, and Hallucination Detection The host calls out the fallacy of setting temperature to zero to stop hallucinations. Mark and Myra explain why supervised probes are superior to unsupervised SAEs for hallucination detection due to feature splitting.29:35–33:42 · Guest disagreement 1/10 Mathematical Equivalence of Activation Steering and In-Context Learning Mark explains team research proving mathematical equivalence between activation steering and in-context learning. The host synthesizes this with KV cache state modifications during inference.33:42–37:21 · Guest disagreement 1/10 Moving Beyond Brute-Force RL Toward Intentional Model Architecture The host compares low-rank adapters to steering. Mark uses the metaphor of modifying the physical pipes versus modifying the water flowing through them to explain intentional architectural design over brute-force RL.37:21–42:25 · Guest disagreement 0/10 Accessibility, Community Infrastructure, and Educational Pathways in MechInterp The hosts and guests discuss the low compute barrier to entry in MechInterp research, educational pathways like MATS, and the launch of dedicated industry tracks at AI Engineer conferences.42:25–46:34 · Guest disagreement 1/10 Code Representations, Steering Limits, and Scaling Frontiers The host notes public skepticism that steering can move beyond stylistic quirks like Gen Z tone into deep reasoning. Myra acknowledges these limits and highlights the need for new learning algorithms beyond simple scale.46:35–52:20 · Guest disagreement 0/10 Interpretability for Life Sciences: Evo 2 and Biomarker Discovery Mark outlines bidirectional interpretability in life sciences models, explaining how Goodfire partnered with Mayo Clinic and Prima Menta to uncover novel biomarkers for Alzheimer's disease.52:20–57:10 · Guest disagreement 1/10 Applying Interpretability to Vision, Diffusion Models, and World Models Myra explains how pixel space offers immediate visual feedback loops for interpretability. Mark discusses how astrophysics models learn alien heuristics that functionally work without grokking standard physical laws.57:10–1:00:53 · Guest disagreement 0/10 Science Fiction Inspirations: Ted Chiang, Model Sandbagging, and Neurobiology The host connects Ted Chiang's story 'Understand' to LLM model sandbagging and superintelligence test-evasion. Mark references Chiang's 'Exhalation' as an allegory for self-interpretability.1:00:53–1:06:14 · Guest disagreement 1/10 Pragmatic AI Safety, Superalignment, and Scalable Oversight The host questions whether weak-to-strong generalization will scale once models surpass human ability. Myra and Mark emphasize pragmatic, grounded alignment through scalable oversight rather than doomerism.2:47–5:40 · The hosts pushing back 0/10 Engineering Backgrounds and Early Team Roles at Goodfire The host and guests share background stories collegially, including overlapping history at Two Sigma and Palantir. The dynamic is conversational and welcoming.5:40–8:08 · The hosts pushing back 1/10 Unpacking Mechanistic Interpretability Beyond Black-Box Methods Co-host Vibu gives a detailed overview of mechanistic interpretability techniques like SAEs and probing. Mark expands the scope from post-hoc poking to intentional architecture design during training.8:09–12:09 · The hosts pushing back 4/10 Post-Training Interventions, Surgical Edits, and Grokking Dynamics The host engages deeply on grokking and double descent, suggesting MechInterp reframes double descent as translating representation space. Mark builds on this by explaining surgical edits and generalizing solutions.12:10–16:58 · The hosts pushing back 3/10 Investigating Subliminal Learning and Shared Latent Geometry When discussing subliminal learning, the host posits a Platonic representation hypothesis. Mark counters that the phenomenon is primarily a path-dependent statistical artifact stemming from shared initial random seeds.16:58–19:04 · The hosts pushing back 1/10 Evaluating Linear Probes Versus Sparse Autoencoders in Practice The host presses on the limitations of Sparse Autoencoders. Myra explains their counterintuitive finding where raw activation probes consistently outperform SAE probes except on noisy datasets.19:04–21:56 · The hosts pushing back 0/10 Enterprise Deployment Case Study: Real-Time Guardrails at Rakuten Vibu notes the runtime latency advantage of internal probes over LLM-as-a-judge guardrails. Mark outlines Rakuten's real-world deployment challenges including synthetic-to-real transfer and token-level PII scrubbing.21:56–25:32 · The hosts pushing back 0/10 Live Demonstration: Real-Time Activation Steering on Kimi K2 Mark showcases a live CLI demonstration steering activations on the 1T parameter Kimi K2 model. The host follows the live output transition into Gen Z slang.25:32–29:35 · The hosts pushing back 1/10 Feature Extraction Pipelines, Labeling, and Hallucination Detection The host calls out the fallacy of setting temperature to zero to stop hallucinations. Mark and Myra explain why supervised probes are superior to unsupervised SAEs for hallucination detection due to feature splitting.29:35–33:42 · The hosts pushing back 3/10 Mathematical Equivalence of Activation Steering and In-Context Learning Mark explains team research proving mathematical equivalence between activation steering and in-context learning. The host synthesizes this with KV cache state modifications during inference.33:42–37:21 · The hosts pushing back 2/10 Moving Beyond Brute-Force RL Toward Intentional Model Architecture The host compares low-rank adapters to steering. Mark uses the metaphor of modifying the physical pipes versus modifying the water flowing through them to explain intentional architectural design over brute-force RL.37:21–42:25 · The hosts pushing back 0/10 Accessibility, Community Infrastructure, and Educational Pathways in MechInterp The hosts and guests discuss the low compute barrier to entry in MechInterp research, educational pathways like MATS, and the launch of dedicated industry tracks at AI Engineer conferences.42:25–46:34 · The hosts pushing back 3/10 Code Representations, Steering Limits, and Scaling Frontiers The host notes public skepticism that steering can move beyond stylistic quirks like Gen Z tone into deep reasoning. Myra acknowledges these limits and highlights the need for new learning algorithms beyond simple scale.46:35–52:20 · The hosts pushing back 0/10 Interpretability for Life Sciences: Evo 2 and Biomarker Discovery Mark outlines bidirectional interpretability in life sciences models, explaining how Goodfire partnered with Mayo Clinic and Prima Menta to uncover novel biomarkers for Alzheimer's disease.52:20–57:10 · The hosts pushing back 2/10 Applying Interpretability to Vision, Diffusion Models, and World Models Myra explains how pixel space offers immediate visual feedback loops for interpretability. Mark discusses how astrophysics models learn alien heuristics that functionally work without grokking standard physical laws.57:10–1:00:53 · The hosts pushing back 1/10 Science Fiction Inspirations: Ted Chiang, Model Sandbagging, and Neurobiology The host connects Ted Chiang's story 'Understand' to LLM model sandbagging and superintelligence test-evasion. Mark references Chiang's 'Exhalation' as an allegory for self-interpretability.1:00:53–1:06:14 · The hosts pushing back 3/10 Pragmatic AI Safety, Superalignment, and Scalable Oversight The host questions whether weak-to-strong generalization will scale once models surpass human ability. Myra and Mark emphasize pragmatic, grounded alignment through scalable oversight rather than doomerism.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 13:29 Rejecting Platonic Representations for Seed Artifacts

Mark directly challenges the host's Platonic representation theory of subliminal learning, clarifying it is a statistical path-dependent artifact from shared initialization seeds.

Hardest push from the hosts ▶ 31:47 Challenging Steering Power Relative to Prompting

The host directly challenges the guest's thesis by asking whether activation steering is fundamentally less powerful than prompting.

Biggest teaching moment ▶ 30:57 Proving Mathematical Equivalence of Steering and ICL

Mark educates the hosts on recent research establishing a quantitative formula mapping activation steering magnitudes directly to in-context learning shots.

The host holds their own ▶ 32:41 Host Maps Steering Directly to KV Cache Updates

The host demonstrates deep technical knowledge by explaining how in-context learning modifies the KV cache and equating its cumulative state to multi-layer activation steering.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Engineering Backgrounds and Early Team Roles at Goodfire 4100 The host and guests share background stories collegially, including overlapping history at Two Sigma and Palantir. The dynamic is conversational and welcoming.
Unpacking Mechanistic Interpretability Beyond Black-Box Methods 7411 Co-host Vibu gives a detailed overview of mechanistic interpretability techniques like SAEs and probing. Mark expands the scope from post-hoc poking to intentional architecture design during training.
Post-Training Interventions, Surgical Edits, and Grokking Dynamics 7514 The host engages deeply on grokking and double descent, suggesting MechInterp reframes double descent as translating representation space. Mark builds on this by explaining surgical edits and generalizing solutions.
Investigating Subliminal Learning and Shared Latent Geometry 6623 When discussing subliminal learning, the host posits a Platonic representation hypothesis. Mark counters that the phenomenon is primarily a path-dependent statistical artifact stemming from shared initial random seeds.
Evaluating Linear Probes Versus Sparse Autoencoders in Practice 5611 The host presses on the limitations of Sparse Autoencoders. Myra explains their counterintuitive finding where raw activation probes consistently outperform SAE probes except on noisy datasets.
Enterprise Deployment Case Study: Real-Time Guardrails at Rakuten 6500 Vibu notes the runtime latency advantage of internal probes over LLM-as-a-judge guardrails. Mark outlines Rakuten's real-world deployment challenges including synthetic-to-real transfer and token-level PII scrubbing.
Live Demonstration: Real-Time Activation Steering on Kimi K2 4500 Mark showcases a live CLI demonstration steering activations on the 1T parameter Kimi K2 model. The host follows the live output transition into Gen Z slang.
Feature Extraction Pipelines, Labeling, and Hallucination Detection 6611 The host calls out the fallacy of setting temperature to zero to stop hallucinations. Mark and Myra explain why supervised probes are superior to unsupervised SAEs for hallucination detection due to feature splitting.
Mathematical Equivalence of Activation Steering and In-Context Learning 7713 Mark explains team research proving mathematical equivalence between activation steering and in-context learning. The host synthesizes this with KV cache state modifications during inference.
Moving Beyond Brute-Force RL Toward Intentional Model Architecture 7612 The host compares low-rank adapters to steering. Mark uses the metaphor of modifying the physical pipes versus modifying the water flowing through them to explain intentional architectural design over brute-force RL.
Accessibility, Community Infrastructure, and Educational Pathways in MechInterp 5400 The hosts and guests discuss the low compute barrier to entry in MechInterp research, educational pathways like MATS, and the launch of dedicated industry tracks at AI Engineer conferences.
Code Representations, Steering Limits, and Scaling Frontiers 6613 The host notes public skepticism that steering can move beyond stylistic quirks like Gen Z tone into deep reasoning. Myra acknowledges these limits and highlights the need for new learning algorithms beyond simple scale.
Interpretability for Life Sciences: Evo 2 and Biomarker Discovery 6700 Mark outlines bidirectional interpretability in life sciences models, explaining how Goodfire partnered with Mayo Clinic and Prima Menta to uncover novel biomarkers for Alzheimer's disease.
Applying Interpretability to Vision, Diffusion Models, and World Models 6612 Myra explains how pixel space offers immediate visual feedback loops for interpretability. Mark discusses how astrophysics models learn alien heuristics that functionally work without grokking standard physical laws.
Science Fiction Inspirations: Ted Chiang, Model Sandbagging, and Neurobiology 7501 The host connects Ted Chiang's story 'Understand' to LLM model sandbagging and superintelligence test-evasion. Mark references Chiang's 'Exhalation' as an allegory for self-interpretability.
Pragmatic AI Safety, Superalignment, and Scalable Oversight 7613 The host questions whether weak-to-strong generalization will scale once models surpass human ability. Myra and Mark emphasize pragmatic, grounded alignment through scalable oversight rather than doomerism.

Statements from this episode (22)

Prediction Not checkable as stated
Deng: Interpretability Will Unlock the Next Frontier of AI Models
“We really believe that interpretability will unlock the new generation, next frontier of safe and powerful AI models.”
Myra Deng Feb 5, 2026 ▶ 1:11
Insight
Bissell: Interpretability is rarely applied during training for model design
“Bring interpretability to training, which I don't think has been done all that much before. A lot of this stuff is sort of post-talk poking at models as opposed to actually using this to intentionally design them.”
Mark Bissell Feb 5, 2026 ▶ 7:58
Assertion Supported
Bissell: CCP bias is identifiable in Qwen and DeepSeek-R1 representation spaces
“Well, there's, there are certainly internal, yeah, parts of the representation space where you can sort of see where that lives.”
Mark Bissell Feb 5, 2026 ▶ 10:08
Opinion
Bissell: Subliminal learning typically only affects models sharing initial random seeds
“I think it only applies to models that were initialized from the same starting Z. Usually, yes.”
Mark Bissell Feb 5, 2026 ▶ 13:24
Disclosure
Deng: Goodfire's first steering API trailed prompting and fine-tuning
“When it comes to like control and design of models, you know, we tried steering with our first API and realized that it still fell short of black box techniques like prompting or fine tuning.”
Myra Deng Feb 5, 2026 ▶ 16:05
Assertion Not checkable as stated
Deng: Raw Activation Probes Often Outperform Sparse Autoencoder Probes
“And we've seen in many cases that probes just trained on raw activations seem to perform better than SAE probes, which is a bit surprising if you think that SAEs are actually also capturing the concepts that you would want to capture cleanly and more surgicall…”
Myra Deng Feb 5, 2026 ▶ 17:35
Assertion Not checkable as stated
Bissell: SAE-Based Approach Proved Most Generalizable for PII Detection
“Although in the PII instance, I think we're into SAE, an SAE based approach actually did prove to be the most generalizable.”
Mark Bissell Feb 5, 2026 ▶ 18:25
Disclosure
Deng: Rakuten Uses Goodfire AI for Production LLM PII Scrubbing
“They are using us to essentially guardrail and inference time monitor their language model usage and their agent usage to detect things like PII so that they don't route private user information to downstream model providers and So that's, you know, going thro…”
Myra Deng Feb 5, 2026 ▶ 19:05
Insight
Bissell: Interpretability Probes Add Virtually No Latency to Model Inference
“And something like a probe is super lightweight. Yeah. It's no extra latency really.”
Mark Bissell Feb 5, 2026 ▶ 21:52
Assertion Supported
Bissell: Goodfire performs activation steering on 1-trillion parameter Kimi K2
“Here you're going to see steering on a one trillion parameter model. This is Kimi K two.”
Mark Bissell Feb 5, 2026 ▶ 22:32
Disclosure
Goodfire is developing interpretability tools to detect model hallucinations
“You really predicted some, a project we're already working on right now, which is detecting hallucinations using interpretability techniques.”
Myra Deng Feb 5, 2026 ▶ 27:30
Assertion Supported
Deng: Models internally represent uncertainty preceding hallucinatory behavior
“We've seen that models internally have some awareness of like uncertainty or some sort of like user pleasing behavior that leads to hallucinatory behavior.”
Myra Deng Feb 5, 2026 ▶ 27:50
Insight
Bissell: Activation Steering and In-Context Learning Are Quantitatively Equivalent
“He actually has a paper that, as well as some, you know, others from the team and elsewhere, that go into the essentially equivalence of activation steering and in-context learning, and how those are from a, he thinks of everything in a cognitive neuroscience …”
Mark Bissell Feb 5, 2026 ▶ 30:57
Assertion Supported
Bissell: Steering Experiments Can Predict Examples Needed for Jailbreaks
“What's in this in context learning and activation steering equivalence paper is you can like predict the number of examples that you will need to put in there in order to jailbreak the model. By doing steering experiments and using this sort of like equivalenc…”
Mark Bissell Feb 5, 2026 ▶ 32:22
Opinion
Bissell: Current AI Training and Post-Training Methods Are Primitive
“I hope that we look back at how we're currently training models and post training models and just think what a primitive way of doing that right now. Like there's no intentionality really in.”
Mark Bissell Feb 5, 2026 ▶ 35:33
Prediction Not checkable as stated
Deng: Scaling alone will not achieve AI needed for mission-critical deployments
“Scale is not going to get us to the type of AI development that we want to be at in, in the future as these models get more powerful and get deployed and all these sorts of like mission critical contexts.”
Myra Deng Feb 5, 2026 ▶ 44:15
Disclosure
Goodfire AI: We replicated code error and malicious features in Llama
“We replicated a lot of these features in, in our llama models as well.”
Myra Deng Feb 5, 2026 ▶ 46:01
Assertion Supported
Goodfire AI applied interpretability with Mayo Clinic to find novel Alzheimer's biomarkers
“We are partnered with Organizations like Mayo Clinic, leading research health system in the United States, our institute, as well as a startup called Prima Menta, which focuses on neurodegenerative disease. And in our partnership with them, we've used foundati…”
Mark Bissell Feb 5, 2026 ▶ 48:43
Insight
Deng: Visual interpretability yields faster feedback cycles than language models
“With language models, when you get features, you still have to do auto interpret and things like that to actually get an understanding of what this concept is. But in image and video and world, it's like extremely easy to grok what the concept is because you c…”
Myra Deng Feb 5, 2026 ▶ 53:36
Assertion Not checkable as stated
Bissell: Physics models learn heuristic shortcuts rather than fundamental rules
“Even if you train certain astrophysics models, it does not learn F equals ma, like the same way that you can, you know, have a model do well for modular arithmetic, but it doesn't really like learn what, how we think of modular arithmetic. It learned some craz…”
Mark Bissell Feb 5, 2026 ▶ 55:55
Insight
Deng: Computational Neuroscientists Are Moving to AI Interpretability for Unfettered Experimental Access
“When we talk to a lot of computational neuroscientists, they Moved to enter because they were like, look, we have unfettered access to this artificial, intelligent mind. It's so much, you have access to everything. You can run as many ablations and experiments…”
Myra Deng Feb 5, 2026 ▶ 1:00:14
Insight
Bissell: AI safety research must scale with superintelligence or fight losing battle
“Ideally, you are setting up your research so that as super intelligence arrives, that is a tailwind. That's also bolstering our ability to like understand the models because otherwise you're fighting a losing battle. If it's like the systems are getting more a…”
Mark Bissell Feb 5, 2026 ▶ 1:04:20
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.