Feb 5, 2026 · 1h 8m · latent-space
Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
Goodfire AI's Mark Bissell and Myra Deng join the Latent Space Podcast to discuss how mechanistic interpretability is evolving from passive model inspection into an active engineering discipline. They demonstrate real-time activation steering, enterprise safety guardrails, and applications spanning life sciences and foundation model architecture design.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Mark directly challenges the host's Platonic representation theory of subliminal learning, clarifying it is a statistical path-dependent artifact from shared initialization seeds.
Hardest push from the hosts ▶ 31:47 Challenging Steering Power Relative to PromptingThe host directly challenges the guest's thesis by asking whether activation steering is fundamentally less powerful than prompting.
Biggest teaching moment ▶ 30:57 Proving Mathematical Equivalence of Steering and ICLMark educates the hosts on recent research establishing a quantitative formula mapping activation steering magnitudes directly to in-context learning shots.
The host holds their own ▶ 32:41 Host Maps Steering Directly to KV Cache UpdatesThe host demonstrates deep technical knowledge by explaining how in-context learning modifies the KV cache and equating its cumulative state to multi-layer activation steering.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Engineering Backgrounds and Early Team Roles at Goodfire | 4 | 1 | 0 | 0 | The host and guests share background stories collegially, including overlapping history at Two Sigma and Palantir. The dynamic is conversational and welcoming. | |
| Unpacking Mechanistic Interpretability Beyond Black-Box Methods | 7 | 4 | 1 | 1 | Co-host Vibu gives a detailed overview of mechanistic interpretability techniques like SAEs and probing. Mark expands the scope from post-hoc poking to intentional architecture design during training. | |
| Post-Training Interventions, Surgical Edits, and Grokking Dynamics | 7 | 5 | 1 | 4 | The host engages deeply on grokking and double descent, suggesting MechInterp reframes double descent as translating representation space. Mark builds on this by explaining surgical edits and generalizing solutions. | |
| Investigating Subliminal Learning and Shared Latent Geometry | 6 | 6 | 2 | 3 | When discussing subliminal learning, the host posits a Platonic representation hypothesis. Mark counters that the phenomenon is primarily a path-dependent statistical artifact stemming from shared initial random seeds. | |
| Evaluating Linear Probes Versus Sparse Autoencoders in Practice | 5 | 6 | 1 | 1 | The host presses on the limitations of Sparse Autoencoders. Myra explains their counterintuitive finding where raw activation probes consistently outperform SAE probes except on noisy datasets. | |
| Enterprise Deployment Case Study: Real-Time Guardrails at Rakuten | 6 | 5 | 0 | 0 | Vibu notes the runtime latency advantage of internal probes over LLM-as-a-judge guardrails. Mark outlines Rakuten's real-world deployment challenges including synthetic-to-real transfer and token-level PII scrubbing. | |
| Live Demonstration: Real-Time Activation Steering on Kimi K2 | 4 | 5 | 0 | 0 | Mark showcases a live CLI demonstration steering activations on the 1T parameter Kimi K2 model. The host follows the live output transition into Gen Z slang. | |
| Feature Extraction Pipelines, Labeling, and Hallucination Detection | 6 | 6 | 1 | 1 | The host calls out the fallacy of setting temperature to zero to stop hallucinations. Mark and Myra explain why supervised probes are superior to unsupervised SAEs for hallucination detection due to feature splitting. | |
| Mathematical Equivalence of Activation Steering and In-Context Learning | 7 | 7 | 1 | 3 | Mark explains team research proving mathematical equivalence between activation steering and in-context learning. The host synthesizes this with KV cache state modifications during inference. | |
| Moving Beyond Brute-Force RL Toward Intentional Model Architecture | 7 | 6 | 1 | 2 | The host compares low-rank adapters to steering. Mark uses the metaphor of modifying the physical pipes versus modifying the water flowing through them to explain intentional architectural design over brute-force RL. | |
| Accessibility, Community Infrastructure, and Educational Pathways in MechInterp | 5 | 4 | 0 | 0 | The hosts and guests discuss the low compute barrier to entry in MechInterp research, educational pathways like MATS, and the launch of dedicated industry tracks at AI Engineer conferences. | |
| Code Representations, Steering Limits, and Scaling Frontiers | 6 | 6 | 1 | 3 | The host notes public skepticism that steering can move beyond stylistic quirks like Gen Z tone into deep reasoning. Myra acknowledges these limits and highlights the need for new learning algorithms beyond simple scale. | |
| Interpretability for Life Sciences: Evo 2 and Biomarker Discovery | 6 | 7 | 0 | 0 | Mark outlines bidirectional interpretability in life sciences models, explaining how Goodfire partnered with Mayo Clinic and Prima Menta to uncover novel biomarkers for Alzheimer's disease. | |
| Applying Interpretability to Vision, Diffusion Models, and World Models | 6 | 6 | 1 | 2 | Myra explains how pixel space offers immediate visual feedback loops for interpretability. Mark discusses how astrophysics models learn alien heuristics that functionally work without grokking standard physical laws. | |
| Science Fiction Inspirations: Ted Chiang, Model Sandbagging, and Neurobiology | 7 | 5 | 0 | 1 | The host connects Ted Chiang's story 'Understand' to LLM model sandbagging and superintelligence test-evasion. Mark references Chiang's 'Exhalation' as an allegory for self-interpretability. | |
| Pragmatic AI Safety, Superalignment, and Scalable Oversight | 7 | 6 | 1 | 3 | The host questions whether weak-to-strong generalization will scale once models surpass human ability. Myra and Mark emphasize pragmatic, grounded alignment through scalable oversight rather than doomerism. |