why aren't all 8 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Supported
Bissell: CCP bias is identifiable in Qwen and DeepSeek-R1 representation spaces
“Well, there's, there are certainly internal, yeah, parts of the representation space where you can sort of see where that lives.”
Assertion Supported
Deng: Models internally represent uncertainty preceding hallucinatory behavior
“We've seen that models internally have some awareness of like uncertainty or some sort of like user pleasing behavior that leads to hallucinatory behavior.”
Assertion Supported
Bissell: Steering Experiments Can Predict Examples Needed for Jailbreaks
“What's in this in context learning and activation steering equivalence paper is you can like predict the number of examples that you will need to put in there in order to jailbreak the model. By doing steering experiments and using this sort of like equivalenc…”
Assertion Supported
Goodfire AI applied interpretability with Mayo Clinic to find novel Alzheimer's biomarkers
“We are partnered with Organizations like Mayo Clinic, leading research health system in the United States, our institute, as well as a startup called Prima Menta, which focuses on neurodegenerative disease. And in our partnership with them, we've used foundati…”
Assertion Supported
Bissell: Goodfire performs activation steering on 1-trillion parameter Kimi K2
“Here you're going to see steering on a one trillion parameter model. This is Kimi K two.”
Assertion Supported
Bissell: Rakuten uses interpretability in production to scrub PII from customer chats
“From Goodfire's perspective, you know, we, so like one of our partners, Rakuten, is deploying an interpretability based tool in production with one of their language agents. This is a really cool use case where if you, What they needed to do was take chats bet…”
Assertion Supported
Bissell: Interpretability allows direct painting into a diffusion model's mental map
“Using interpretability techniques, you can sort of like plug directly into the mind of the model, and you get a two D canvas where you can basically like paint directly into its mental map of the image. And so we used unsupervised techniques to basically figur…”
Assertion Supported
Mechanistic Interpretability Scales Without Bottlenecks to Large Models Like DeepSeek
“There's no gap for scale. Like, they've shown that even for the biggest open source models, you like, even like DeepSeq's big models, you, they can do it. And then in general, like, scaling is not the bottleneck.”