why aren't all 21 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Supported
Bissell: CCP bias is identifiable in Qwen and DeepSeek-R1 representation spaces
“Well, there's, there are certainly internal, yeah, parts of the representation space where you can sort of see where that lives.”
Prediction Not checkable as stated
Deng: Scaling alone will not achieve AI needed for mission-critical deployments
“Scale is not going to get us to the type of AI development that we want to be at in, in the future as these models get more powerful and get deployed and all these sorts of like mission critical contexts.”
Assertion Supported
Deng: Models internally represent uncertainty preceding hallucinatory behavior
“We've seen that models internally have some awareness of like uncertainty or some sort of like user pleasing behavior that leads to hallucinatory behavior.”
Insight
Bissell: Activation Steering and In-Context Learning Are Quantitatively Equivalent
“He actually has a paper that, as well as some, you know, others from the team and elsewhere, that go into the essentially equivalence of activation steering and in-context learning, and how those are from a, he thinks of everything in a cognitive neuroscience …”
Assertion Not checkable as stated
Deng: Raw Activation Probes Often Outperform Sparse Autoencoder Probes
“And we've seen in many cases that probes just trained on raw activations seem to perform better than SAE probes, which is a bit surprising if you think that SAEs are actually also capturing the concepts that you would want to capture cleanly and more surgicall…”
Assertion Supported
Bissell: Steering Experiments Can Predict Examples Needed for Jailbreaks
“What's in this in context learning and activation steering equivalence paper is you can like predict the number of examples that you will need to put in there in order to jailbreak the model. By doing steering experiments and using this sort of like equivalenc…”
Assertion Supported
Goodfire AI applied interpretability with Mayo Clinic to find novel Alzheimer's biomarkers
“We are partnered with Organizations like Mayo Clinic, leading research health system in the United States, our institute, as well as a startup called Prima Menta, which focuses on neurodegenerative disease. And in our partnership with them, we've used foundati…”
Assertion Supported
Bissell: Goodfire performs activation steering on 1-trillion parameter Kimi K2
“Here you're going to see steering on a one trillion parameter model. This is Kimi K two.”
Assertion Supported
Bissell: Rakuten uses interpretability in production to scrub PII from customer chats
“From Goodfire's perspective, you know, we, so like one of our partners, Rakuten, is deploying an interpretability based tool in production with one of their language agents. This is a really cool use case where if you, What they needed to do was take chats bet…”
Assertion Not checkable as stated
Bissell: SAE-Based Approach Proved Most Generalizable for PII Detection
“Although in the PII instance, I think we're into SAE, an SAE based approach actually did prove to be the most generalizable.”
Insight
Deng: Visual interpretability yields faster feedback cycles than language models
“With language models, when you get features, you still have to do auto interpret and things like that to actually get an understanding of what this concept is. But in image and video and world, it's like extremely easy to grok what the concept is because you c…”
Disclosure
Deng: Goodfire's first steering API trailed prompting and fine-tuning
“When it comes to like control and design of models, you know, we tried steering with our first API and realized that it still fell short of black box techniques like prompting or fine tuning.”
Insight
Bissell: Interpretability is rarely applied during training for model design
“Bring interpretability to training, which I don't think has been done all that much before. A lot of this stuff is sort of post-talk poking at models as opposed to actually using this to intentionally design them.”
Prediction Not checkable as stated
Deng: Interpretability Will Unlock the Next Frontier of AI Models
“We really believe that interpretability will unlock the new generation, next frontier of safe and powerful AI models.”
Insight
Merullo: LLM memorization spans a gradient from reasoning to rote recall
“You can actually see, like the way that we, like, disentangle memorization, you can kind of see this like, gradient of memorization in between both mechanistically and behaviorally with, like, logical reasoning tasks being quite distinct from rote memorization…”
Insight
Bissell: Mechanistic interpretability provides power-user tools for manipulating AI models
“Interpretability gives you a set of, I think of it almost as like power user tools for accessing models and doing things with them that you might not have realized you could.”
Assertion Supported
Bissell: Interpretability allows direct painting into a diffusion model's mental map
“Using interpretability techniques, you can sort of like plug directly into the mind of the model, and you get a two D canvas where you can basically like paint directly into its mental map of the image. And so we used unsupervised techniques to basically figur…”
Assertion Supported
Mechanistic Interpretability Scales Without Bottlenecks to Large Models Like DeepSeek
“There's no gap for scale. Like, they've shown that even for the biggest open source models, you like, even like DeepSeq's big models, you, they can do it. And then in general, like, scaling is not the bottleneck.”
Disclosure
Goodfire is developing interpretability tools to detect model hallucinations
“You really predicted some, a project we're already working on right now, which is detecting hallucinations using interpretability techniques.”
Disclosure
Deng: Rakuten Uses Goodfire AI for Production LLM PII Scrubbing
“They are using us to essentially guardrail and inference time monitor their language model usage and their agent usage to detect things like PII so that they don't route private user information to downstream model providers and So that's, you know, going thro…”
Disclosure
Goodfire AI: We replicated code error and malicious features in Llama
“We replicated a lot of these features in, in our llama models as well.”