“We've seen that models internally have some awareness of like uncertainty or some sort of like user pleasing behavior that leads to hallucinatory behavior.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Myra Deng
PredictionNot checkable as stated
Deng: Scaling alone will not achieve AI needed for mission-critical deployments
“Scale is not going to get us to the type of AI development that we want to be at in, in the future as these models get more powerful and get deployed and all these sorts of like mission critical contexts.”
Myra DengFeb 5, 2026▶ 44:15Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
AssertionNot checkable as stated
Deng: Raw Activation Probes Often Outperform Sparse Autoencoder Probes
“And we've seen in many cases that probes just trained on raw activations seem to perform better than SAE probes, which is a bit surprising if you think that SAEs are actually also capturing the concepts that you would want to capture cleanly and more surgicall…”
Myra DengFeb 5, 2026▶ 17:35Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
PredictionNot checkable as stated
Deng: Interpretability Will Unlock the Next Frontier of AI Models
“We really believe that interpretability will unlock the new generation, next frontier of safe and powerful AI models.”
Myra DengFeb 5, 2026▶ 1:11Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Disclosure
Deng: Goodfire's first steering API trailed prompting and fine-tuning
“When it comes to like control and design of models, you know, we tried steering with our first API and realized that it still fell short of black box techniques like prompting or fine tuning.”
Myra DengFeb 5, 2026▶ 16:05Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Insight
Deng: Visual interpretability yields faster feedback cycles than language models
“With language models, when you get features, you still have to do auto interpret and things like that to actually get an understanding of what this concept is. But in image and video and world, it's like extremely easy to grok what the concept is because you c…”
Myra DengFeb 5, 2026▶ 53:36Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Disclosure
Deng: Rakuten Uses Goodfire AI for Production LLM PII Scrubbing
“They are using us to essentially guardrail and inference time monitor their language model usage and their agent usage to detect things like PII so that they don't route private user information to downstream model providers and So that's, you know, going thro…”
Myra DengFeb 5, 2026▶ 19:05Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 200 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.