Assertion Supported AI assessment confidence: 95% certainty 4/5 debate potential 4/5

Bissell: CCP bias is identifiable in Qwen and DeepSeek-R1 representation spaces

Mark Bissell · Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell · Feb 5, 2026 · at 10:08

Mark Bissell, Member of Technical Staff at Goodfire AI, discusses using mechanistic interpretability to identify and surgically extract political bias vectors from Chinese foundation models like Qwen and DeepSeek-R1.

0:00 / 0:06exact quote · 6.9s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“Well, there's, there are certainly internal, yeah, parts of the representation space where you can sort of see where that lives.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Mark Bissell

Insight
Bissell: Activation Steering and In-Context Learning Are Quantitatively Equivalent
“He actually has a paper that, as well as some, you know, others from the team and elsewhere, that go into the essentially equivalence of activation steering and in-context learning, and how those are from a, he thinks of everything in a cognitive neuroscience …”
Mark Bissell Feb 5, 2026 ▶ 30:57 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Insight
Bissell: Probing internal model features matches LLM-as-a-judge quality at 500x lower cost
“If you ask that model, try to like use it as an LLM as a judge, it's not very good. But if you probe its mind and you sort of detect when the features related to personally identifiable information are firing, that gets you the highest recall of anything. It's…”
Mark Bissell Dec 31, 2025 ▶ 9:53 [State of MechInterp] SAEs in Production, Circuit Tracing, AI4Science, "Pragmatic" Interp — Goodfire
Opinion
Bissell: Subliminal learning typically only affects models sharing initial random seeds
“I think it only applies to models that were initialized from the same starting Z. Usually, yes.”
Mark Bissell Feb 5, 2026 ▶ 13:24 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Assertion Supported
Bissell: Goodfire performs activation steering on 1-trillion parameter Kimi K2
“Here you're going to see steering on a one trillion parameter model. This is Kimi K two.”
Mark Bissell Feb 5, 2026 ▶ 22:32 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Assertion Supported
Bissell: Steering Experiments Can Predict Examples Needed for Jailbreaks
“What's in this in context learning and activation steering equivalence paper is you can like predict the number of examples that you will need to put in there in order to jailbreak the model. By doing steering experiments and using this sort of like equivalenc…”
Mark Bissell Feb 5, 2026 ▶ 32:22 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Opinion
Bissell: Current AI Training and Post-Training Methods Are Primitive
“I hope that we look back at how we're currently training models and post training models and just think what a primitive way of doing that right now. Like there's no intentionality really in.”
Mark Bissell Feb 5, 2026 ▶ 35:33 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.