Assertion Supported
Bissell: CCP bias is identifiable in Qwen and DeepSeek-R1 representation spaces
“Well, there's, there are certainly internal, yeah, parts of the representation space where you can sort of see where that lives.”
Insight
Bissell: Activation Steering and In-Context Learning Are Quantitatively Equivalent
“He actually has a paper that, as well as some, you know, others from the team and elsewhere, that go into the essentially equivalence of activation steering and in-context learning, and how those are from a, he thinks of everything in a cognitive neuroscience …”
Insight
Bissell: Probing internal model features matches LLM-as-a-judge quality at 500x lower cost
“If you ask that model, try to like use it as an LLM as a judge, it's not very good. But if you probe its mind and you sort of detect when the features related to personally identifiable information are firing, that gets you the highest recall of anything. It's…”
Opinion
Bissell: Subliminal learning typically only affects models sharing initial random seeds
“I think it only applies to models that were initialized from the same starting Z. Usually, yes.”
Assertion Supported
Bissell: Goodfire performs activation steering on 1-trillion parameter Kimi K2
“Here you're going to see steering on a one trillion parameter model. This is Kimi K two.”
Assertion Supported
Bissell: Steering Experiments Can Predict Examples Needed for Jailbreaks
“What's in this in context learning and activation steering equivalence paper is you can like predict the number of examples that you will need to put in there in order to jailbreak the model. By doing steering experiments and using this sort of like equivalenc…”
Opinion
Bissell: Current AI Training and Post-Training Methods Are Primitive
“I hope that we look back at how we're currently training models and post training models and just think what a primitive way of doing that right now. Like there's no intentionality really in.”
Assertion Supported
Goodfire AI applied interpretability with Mayo Clinic to find novel Alzheimer's biomarkers
“We are partnered with Organizations like Mayo Clinic, leading research health system in the United States, our institute, as well as a startup called Prima Menta, which focuses on neurodegenerative disease. And in our partnership with them, we've used foundati…”
Assertion Supported
Bissell: Rakuten uses interpretability in production to scrub PII from customer chats
“From Goodfire's perspective, you know, we, so like one of our partners, Rakuten, is deploying an interpretability based tool in production with one of their language agents. This is a really cool use case where if you, What they needed to do was take chats bet…”
Insight
Bissell: Interpretability is rarely applied during training for model design
“Bring interpretability to training, which I don't think has been done all that much before. A lot of this stuff is sort of post-talk poking at models as opposed to actually using this to intentionally design them.”
Assertion Not checkable as stated
Bissell: SAE-Based Approach Proved Most Generalizable for PII Detection
“Although in the PII instance, I think we're into SAE, an SAE based approach actually did prove to be the most generalizable.”
Insight
Bissell: Interpretability Probes Add Virtually No Latency to Model Inference
“And something like a probe is super lightweight. Yeah. It's no extra latency really.”
Assertion Not checkable as stated
Bissell: Physics models learn heuristic shortcuts rather than fundamental rules
“Even if you train certain astrophysics models, it does not learn F equals ma, like the same way that you can, you know, have a model do well for modular arithmetic, but it doesn't really like learn what, how we think of modular arithmetic. It learned some craz…”
Insight
Bissell: AI safety research must scale with superintelligence or fight losing battle
“Ideally, you are setting up your research so that as super intelligence arrives, that is a tailwind. That's also bolstering our ability to like understand the models because otherwise you're fighting a losing battle. If it's like the systems are getting more a…”
Insight
Bissell: Mechanistic interpretability provides power-user tools for manipulating AI models
“Interpretability gives you a set of, I think of it almost as like power user tools for accessing models and doing things with them that you might not have realized you could.”
Assertion Supported
Bissell: Interpretability allows direct painting into a diffusion model's mental map
“Using interpretability techniques, you can sort of like plug directly into the mind of the model, and you get a two D canvas where you can basically like paint directly into its mental map of the image. And so we used unsupervised techniques to basically figur…”