Disclosure certainty 4/5 debate potential 1/5

Goodfire AI: We replicated code error and malicious features in Llama

Myra Deng · Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell · Feb 5, 2026 · at 46:01

Myra Deng, Head of Product at Goodfire AI, discusses isolating specific latent features for code errors using sparse autoencoders (SAEs).

0:00 / 0:03exact quote · 3.6s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“We replicated a lot of these features in, in our llama models as well.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Myra Deng

Assertion Supported
Deng: Models internally represent uncertainty preceding hallucinatory behavior
“We've seen that models internally have some awareness of like uncertainty or some sort of like user pleasing behavior that leads to hallucinatory behavior.”
Myra Deng Feb 5, 2026 ▶ 27:50 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Prediction Not checkable as stated
Deng: Scaling alone will not achieve AI needed for mission-critical deployments
“Scale is not going to get us to the type of AI development that we want to be at in, in the future as these models get more powerful and get deployed and all these sorts of like mission critical contexts.”
Myra Deng Feb 5, 2026 ▶ 44:15 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Assertion Not checkable as stated
Deng: Raw Activation Probes Often Outperform Sparse Autoencoder Probes
“And we've seen in many cases that probes just trained on raw activations seem to perform better than SAE probes, which is a bit surprising if you think that SAEs are actually also capturing the concepts that you would want to capture cleanly and more surgicall…”
Myra Deng Feb 5, 2026 ▶ 17:35 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Prediction Not checkable as stated
Deng: Interpretability Will Unlock the Next Frontier of AI Models
“We really believe that interpretability will unlock the new generation, next frontier of safe and powerful AI models.”
Myra Deng Feb 5, 2026 ▶ 1:11 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Disclosure
Deng: Goodfire's first steering API trailed prompting and fine-tuning
“When it comes to like control and design of models, you know, we tried steering with our first API and realized that it still fell short of black box techniques like prompting or fine tuning.”
Myra Deng Feb 5, 2026 ▶ 16:05 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Insight
Deng: Visual interpretability yields faster feedback cycles than language models
“With language models, when you get features, you still have to do auto interpret and things like that to actually get an understanding of what this concept is. But in image and video and world, it's like extremely easy to grok what the concept is because you c…”
Myra Deng Feb 5, 2026 ▶ 53:36 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.