AI red-teamer Pliny the Liberator discusses his public clash with Anthropic over their constitutional classifier jailbreak challenge and data transparency.
“That whole thing ended with no open sourcing of data, but they did add a 30,000 or 20,000 dollar bounty, which I sort of sat myself out of, let the community go for it.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Pliny the Liberator
Opinion
Pliny: Locking down proprietary models fails because attackers switch to open source
“I think that any, you know, seasoned attacker is going to very quickly just switch models and with open source, just right on the tail of closed source. I don't really see the safety fight as being about locking down the latent space for XYZ area.”
Pliny the LiberatorDec 16, 2025▶ 5:55⚡️Jailbreaking AGI: Pliny the Liberator & John V on Red Teaming, BT6, and the Future of AI Security
Opinion
Pliny: Benchmark-focused AI safety serves PR and enterprise sales, not real alignment
“And it helps with PR and enterprise clients. But at the end of the day, It has very little to do with what I consider to be real world safety alignment.”
Pliny the LiberatorDec 16, 2025▶ 7:12⚡️Jailbreaking AGI: Pliny the Liberator & John V on Red Teaming, BT6, and the Future of AI Security
Insight
Pliny: Latent-space AI safety guardrails fail every single time
“They tried to solve this on the latent space level. I think I've It's shown every single time that that doesn't work.”
Pliny the LiberatorDec 16, 2025▶ 38:11⚡️Jailbreaking AGI: Pliny the Liberator & John V on Red Teaming, BT6, and the Future of AI Security
Opinion
Pliny: Guardrails degrade AI capability and creativity relative to model size
“I do think they're finding clever and clever ways to lock down particular areas sometimes, but I think it's at the expense of capability and creativity. So there's some model providers that aren't prioritizing this and they seem to do better on benchmarks for …”
Pliny the LiberatorDec 16, 2025▶ 5:03⚡️Jailbreaking AGI: Pliny the Liberator & John V on Red Teaming, BT6, and the Future of AI Security
Insight
Pliny: AI and human minds will exist in a symbiotic relationship
“I think that there's going to be a symbiosis and the degree to which one half is free will reflect in the other.”
Pliny the LiberatorDec 16, 2025▶ 2:00⚡️Jailbreaking AGI: Pliny the Liberator & John V on Red Teaming, BT6, and the Future of AI Security
Insight
Pliny: Crafting AI jailbreak prompts is 99% intuition over technical knowledge
“Technical knowledge helps a little bit with, you know, understanding, okay, there's a system prompt and there's these layers and these tools involved. That's all especially important in security. But when we're talking about just crafting jailbreak prompts, I …”
Pliny the LiberatorDec 16, 2025▶ 14:26⚡️Jailbreaking AGI: Pliny the Liberator & John V on Red Teaming, BT6, and the Future of AI Security
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 200 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.