Myra Deng

Head of Product, Goodfire AI · 1 appearance on the record.

computed by AI from the episodes · how this works → · full disclaimer →

executiveoperatorengineer@myra_deng ↗LinkedIn ↗myradeng.com ↗

Myra Deng is the Head of Product at Goodfire AI, where she leads development of tools and APIs for mechanistic interpretability. Previously, she worked as a quantitative engineer and product lead at Two Sigma.

10statements → 4claims → 1claims resolved → 3.9/5average certainty → 1.9/5average debate potential →

1 supported 0 partly supported 0 contradicted 3 not checkable as stated how the 4 claims stand · each chip opens the sources

2 predictions · 2 assertions · 2 insights · 4 disclosures · every statement was checked. The predictions and assertions are the 4 claims: statements the public record can support or contradict. 1 is resolved, and 3 name no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Myra argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
Deng: Models internally represent uncertainty preceding hallucinatory behavior
“We've seen that models internally have some awareness of like uncertainty or some sort of like user pleasing behavior that leads to hallucinatory behavior.”
Myra Deng Feb 5, 2026 ▶ 27:50 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell

Everything Myra Deng said on Latent Space that made the record, most notable first. Filter by type, assessment or year in the ledger →

Assertion Supported
Deng: Models internally represent uncertainty preceding hallucinatory behavior
“We've seen that models internally have some awareness of like uncertainty or some sort of like user pleasing behavior that leads to hallucinatory behavior.”
Myra Deng Feb 5, 2026 ▶ 27:50 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Prediction Not checkable as stated
Deng: Scaling alone will not achieve AI needed for mission-critical deployments
“Scale is not going to get us to the type of AI development that we want to be at in, in the future as these models get more powerful and get deployed and all these sorts of like mission critical contexts.”
Myra Deng Feb 5, 2026 ▶ 44:15 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Assertion Not checkable as stated
Deng: Raw Activation Probes Often Outperform Sparse Autoencoder Probes
“And we've seen in many cases that probes just trained on raw activations seem to perform better than SAE probes, which is a bit surprising if you think that SAEs are actually also capturing the concepts that you would want to capture cleanly and more surgicall…”
Myra Deng Feb 5, 2026 ▶ 17:35 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Prediction Not checkable as stated
Deng: Interpretability Will Unlock the Next Frontier of AI Models
“We really believe that interpretability will unlock the new generation, next frontier of safe and powerful AI models.”
Myra Deng Feb 5, 2026 ▶ 1:11 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Disclosure
Deng: Goodfire's first steering API trailed prompting and fine-tuning
“When it comes to like control and design of models, you know, we tried steering with our first API and realized that it still fell short of black box techniques like prompting or fine tuning.”
Myra Deng Feb 5, 2026 ▶ 16:05 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Insight
Deng: Visual interpretability yields faster feedback cycles than language models
“With language models, when you get features, you still have to do auto interpret and things like that to actually get an understanding of what this concept is. But in image and video and world, it's like extremely easy to grok what the concept is because you c…”
Myra Deng Feb 5, 2026 ▶ 53:36 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Disclosure
Deng: Rakuten Uses Goodfire AI for Production LLM PII Scrubbing
“They are using us to essentially guardrail and inference time monitor their language model usage and their agent usage to detect things like PII so that they don't route private user information to downstream model providers and So that's, you know, going thro…”
Myra Deng Feb 5, 2026 ▶ 19:05 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Disclosure
Goodfire is developing interpretability tools to detect model hallucinations
“You really predicted some, a project we're already working on right now, which is detecting hallucinations using interpretability techniques.”
Myra Deng Feb 5, 2026 ▶ 27:30 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Disclosure
Goodfire AI: We replicated code error and malicious features in Llama
“We replicated a lot of these features in, in our llama models as well.”
Myra Deng Feb 5, 2026 ▶ 46:01 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Insight
Deng: Computational Neuroscientists Are Moving to AI Interpretability for Unfettered Experimental Access
“When we talk to a lot of computational neuroscientists, they Moved to enter because they were like, look, we have unfettered access to this artificial, intelligent mind. It's so much, you have access to everything. You can run as many ablations and experiments…”
Myra Deng Feb 5, 2026 ▶ 1:00:14 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell

Appearances (1)

EpisodeDateSpeaking time
Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mar Feb 5, 2026 16m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.