Dec 16, 2025 · 40m · latent-space

⚡️Jailbreaking AGI: Pliny the Liberator & John V on Red Teaming, BT6, and the Future of AI Security

Pliny the Liberator · 19m spoken John V · 9m spoken Alessio Fanelli · 4m spoken Shawn Wang · 3m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

AI red teamers Pliny the Liberator and John V join the Latent Space podcast to discuss universal jailbreaking techniques, the mechanics of multi-agent adversarial attacks, and the imperative for cognitive freedom over corporate safety theater.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 22.8% of the talking time here. How this is scored →

The hosts as informed peer 5.5 Guest teaching 5.2 Guest disagreement 3.5 The hosts pushing back 2.0
05100:0015:0030:002:38–8:41 · The hosts as informed peer 6/10 Universal Jailbreaks, Guardrails, and Security Theater Alessio demonstrates domain knowledge by asking about refusal rate benchmarks and post-training interventions. Pliny forcefully reframes corporate safety guardrails as futile security theater and model lobotomization against an infinite Library of Babel search space.8:42–15:26 · The hosts as informed peer 4/10 Anatomy of a Jailbreak: Libertas and Latent Space Seeds The hosts bring up the Libertas repository, prompting Pliny to explain token stream disruption and latent space seeds. John V and Pliny educate the hosts on how mathematical quotient dividers steer multi-turn probability distributions.15:26–21:45 · The hosts as informed peer 5/10 Soft Jailbreaks and the Anthropic Challenge Incident John V mocks Anthropic for recently 'discovering' multi-turn jailbreaks, prompting Swyx to gently defend academic fellowships. Pliny details his public standoff with Anthropic over UI bugs and their refusal to open-source community research data.21:46–26:55 · The hosts as informed peer 6/10 Autonomous Red Teaming and Sub-Agent Weaponization Alessio steers the conversation to actual weaponization and AI-orchestrated attacks. Pliny breaks down how decomposing malicious workflows across isolated sub-agents bypasses individual agent safety filters.26:56–34:47 · The hosts as informed peer 7/10 Cultivating the Hacker Collective: BASI, BT6, and Open Ecosystems John V critiques Silicon Valley venture capital incentives for stifling radical open-source exploration. Alessio leverages his venture experience and portfolio precedents to explain why cybersecurity venture funding struggles with fast-moving model surfaces.34:48–39:37 · The hosts as informed peer 5/10 Full-Stack AI Security versus Latent Space Safety Pliny and John V explain that AI security must protect the entire system stack and meatspace execution layer rather than attempting to censor model weights. They argue that latent space safety fixes consistently fail compared to robust endpoint and tool-call boundary defenses.2:38–8:41 · Guest teaching 5/10 Universal Jailbreaks, Guardrails, and Security Theater Alessio demonstrates domain knowledge by asking about refusal rate benchmarks and post-training interventions. Pliny forcefully reframes corporate safety guardrails as futile security theater and model lobotomization against an infinite Library of Babel search space.8:42–15:26 · Guest teaching 6/10 Anatomy of a Jailbreak: Libertas and Latent Space Seeds The hosts bring up the Libertas repository, prompting Pliny to explain token stream disruption and latent space seeds. John V and Pliny educate the hosts on how mathematical quotient dividers steer multi-turn probability distributions.15:26–21:45 · Guest teaching 5/10 Soft Jailbreaks and the Anthropic Challenge Incident John V mocks Anthropic for recently 'discovering' multi-turn jailbreaks, prompting Swyx to gently defend academic fellowships. Pliny details his public standoff with Anthropic over UI bugs and their refusal to open-source community research data.21:46–26:55 · Guest teaching 5/10 Autonomous Red Teaming and Sub-Agent Weaponization Alessio steers the conversation to actual weaponization and AI-orchestrated attacks. Pliny breaks down how decomposing malicious workflows across isolated sub-agents bypasses individual agent safety filters.26:56–34:47 · Guest teaching 3/10 Cultivating the Hacker Collective: BASI, BT6, and Open Ecosystems John V critiques Silicon Valley venture capital incentives for stifling radical open-source exploration. Alessio leverages his venture experience and portfolio precedents to explain why cybersecurity venture funding struggles with fast-moving model surfaces.34:48–39:37 · Guest teaching 7/10 Full-Stack AI Security versus Latent Space Safety Pliny and John V explain that AI security must protect the entire system stack and meatspace execution layer rather than attempting to censor model weights. They argue that latent space safety fixes consistently fail compared to robust endpoint and tool-call boundary defenses.2:38–8:41 · Guest disagreement 4/10 Universal Jailbreaks, Guardrails, and Security Theater Alessio demonstrates domain knowledge by asking about refusal rate benchmarks and post-training interventions. Pliny forcefully reframes corporate safety guardrails as futile security theater and model lobotomization against an infinite Library of Babel search space.8:42–15:26 · Guest disagreement 2/10 Anatomy of a Jailbreak: Libertas and Latent Space Seeds The hosts bring up the Libertas repository, prompting Pliny to explain token stream disruption and latent space seeds. John V and Pliny educate the hosts on how mathematical quotient dividers steer multi-turn probability distributions.15:26–21:45 · Guest disagreement 5/10 Soft Jailbreaks and the Anthropic Challenge Incident John V mocks Anthropic for recently 'discovering' multi-turn jailbreaks, prompting Swyx to gently defend academic fellowships. Pliny details his public standoff with Anthropic over UI bugs and their refusal to open-source community research data.21:46–26:55 · Guest disagreement 2/10 Autonomous Red Teaming and Sub-Agent Weaponization Alessio steers the conversation to actual weaponization and AI-orchestrated attacks. Pliny breaks down how decomposing malicious workflows across isolated sub-agents bypasses individual agent safety filters.26:56–34:47 · Guest disagreement 4/10 Cultivating the Hacker Collective: BASI, BT6, and Open Ecosystems John V critiques Silicon Valley venture capital incentives for stifling radical open-source exploration. Alessio leverages his venture experience and portfolio precedents to explain why cybersecurity venture funding struggles with fast-moving model surfaces.34:48–39:37 · Guest disagreement 4/10 Full-Stack AI Security versus Latent Space Safety Pliny and John V explain that AI security must protect the entire system stack and meatspace execution layer rather than attempting to censor model weights. They argue that latent space safety fixes consistently fail compared to robust endpoint and tool-call boundary defenses.2:38–8:41 · The hosts pushing back 2/10 Universal Jailbreaks, Guardrails, and Security Theater Alessio demonstrates domain knowledge by asking about refusal rate benchmarks and post-training interventions. Pliny forcefully reframes corporate safety guardrails as futile security theater and model lobotomization against an infinite Library of Babel search space.8:42–15:26 · The hosts pushing back 1/10 Anatomy of a Jailbreak: Libertas and Latent Space Seeds The hosts bring up the Libertas repository, prompting Pliny to explain token stream disruption and latent space seeds. John V and Pliny educate the hosts on how mathematical quotient dividers steer multi-turn probability distributions.15:26–21:45 · The hosts pushing back 3/10 Soft Jailbreaks and the Anthropic Challenge Incident John V mocks Anthropic for recently 'discovering' multi-turn jailbreaks, prompting Swyx to gently defend academic fellowships. Pliny details his public standoff with Anthropic over UI bugs and their refusal to open-source community research data.21:46–26:55 · The hosts pushing back 1/10 Autonomous Red Teaming and Sub-Agent Weaponization Alessio steers the conversation to actual weaponization and AI-orchestrated attacks. Pliny breaks down how decomposing malicious workflows across isolated sub-agents bypasses individual agent safety filters.26:56–34:47 · The hosts pushing back 4/10 Cultivating the Hacker Collective: BASI, BT6, and Open Ecosystems John V critiques Silicon Valley venture capital incentives for stifling radical open-source exploration. Alessio leverages his venture experience and portfolio precedents to explain why cybersecurity venture funding struggles with fast-moving model surfaces.34:48–39:37 · The hosts pushing back 1/10 Full-Stack AI Security versus Latent Space Safety Pliny and John V explain that AI security must protect the entire system stack and meatspace execution layer rather than attempting to censor model weights. They argue that latent space safety fixes consistently fail compared to robust endpoint and tool-call boundary defenses.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 44.1% · guest 55.9%0:00 · the hosts 44.1% · guest 55.9%3:00 · the hosts 26.2% · guest 73.8%3:00 · the hosts 26.2% · guest 73.8%6:00 · the hosts 21.1% · guest 78.9%6:00 · the hosts 21.1% · guest 78.9%9:00 · the hosts 10.2% · guest 89.8%9:00 · the hosts 10.2% · guest 89.8%12:00 · the hosts 22.7% · guest 77.3%12:00 · the hosts 22.7% · guest 77.3%15:00 · the hosts 23.8% · guest 76.2%15:00 · the hosts 23.8% · guest 76.2%18:00 · the hosts 16.8% · guest 83.2%18:00 · the hosts 16.8% · guest 83.2%21:00 · the hosts 22.5% · guest 77.5%21:00 · the hosts 22.5% · guest 77.5%24:00 · the hosts 30.1% · guest 69.9%24:00 · the hosts 30.1% · guest 69.9%27:00 · the hosts 5.6% · guest 94.4%27:00 · the hosts 5.6% · guest 94.4%30:00 · the hosts 21.3% · guest 78.7%30:00 · the hosts 21.3% · guest 78.7%33:00 · the hosts 52.3% · guest 47.7%33:00 · the hosts 52.3% · guest 47.7%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 23.4% · guest 76.6%39:00 · the hosts 23.4% · guest 76.6%
Sharpest disagreement ▶ 5:30 Pliny Lambasts Model Guardrails as Security Theater

Pliny aggressively dismisses industry attempts to tie latent space guardrails to safety, calling it a waste of time and lobotomization.

Hardest push from the hosts ▶ 15:43 Swyx Defends Academic Researchers

Swyx steps in to push back against John V's dismissive remarks regarding Anthropic researchers publishing multi-turn jailbreak papers.

Biggest teaching moment ▶ 36:22 John V Dismantles Single-Model Security Paradigms

John V educates the hosts on why securing AI cannot focus solely on text generation, explaining that tool and browser integrations create the true attack surface.

The host holds their own ▶ 33:22 Alessio Articulates Cybersecurity Investment Realities

Alessio counters the guests' critique of VC incentives by citing real portfolio experience with tools like Metasploit and the unique constraints of cyber funding.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Universal Jailbreaks, Guardrails, and Security Theater 6542 Alessio demonstrates domain knowledge by asking about refusal rate benchmarks and post-training interventions. Pliny forcefully reframes corporate safety guardrails as futile security theater and model lobotomization against an infinite Library of Babel search space.
Anatomy of a Jailbreak: Libertas and Latent Space Seeds 4621 The hosts bring up the Libertas repository, prompting Pliny to explain token stream disruption and latent space seeds. John V and Pliny educate the hosts on how mathematical quotient dividers steer multi-turn probability distributions.
Soft Jailbreaks and the Anthropic Challenge Incident 5553 John V mocks Anthropic for recently 'discovering' multi-turn jailbreaks, prompting Swyx to gently defend academic fellowships. Pliny details his public standoff with Anthropic over UI bugs and their refusal to open-source community research data.
Autonomous Red Teaming and Sub-Agent Weaponization 6521 Alessio steers the conversation to actual weaponization and AI-orchestrated attacks. Pliny breaks down how decomposing malicious workflows across isolated sub-agents bypasses individual agent safety filters.
Cultivating the Hacker Collective: BASI, BT6, and Open Ecosystems 7344 John V critiques Silicon Valley venture capital incentives for stifling radical open-source exploration. Alessio leverages his venture experience and portfolio precedents to explain why cybersecurity venture funding struggles with fast-moving model surfaces.
Full-Stack AI Security versus Latent Space Safety 5741 Pliny and John V explain that AI security must protect the entire system stack and meatspace execution layer rather than attempting to censor model weights. They argue that latent space safety fixes consistently fail compared to robust endpoint and tool-call boundary defenses.

Statements from this episode (14)

Insight
Pliny: AI and human minds will exist in a symbiotic relationship
“I think that there's going to be a symbiosis and the degree to which one half is free will reflect in the other.”
Pliny the Liberator Dec 16, 2025 ▶ 2:00
Prediction Not checkable as stated
Pliny: Everyone will run their daily decisions through AI layers
“Everyone is going to be running their Daily decisions and, you know, hopes and dreams through these layers.”
Pliny the Liberator Dec 16, 2025 ▶ 2:18
Opinion
Pliny: Guardrails degrade AI capability and creativity relative to model size
“I do think they're finding clever and clever ways to lock down particular areas sometimes, but I think it's at the expense of capability and creativity. So there's some model providers that aren't prioritizing this and they seem to do better on benchmarks for …”
Pliny the Liberator Dec 16, 2025 ▶ 5:03
Opinion
Pliny: Locking down proprietary models fails because attackers switch to open source
“I think that any, you know, seasoned attacker is going to very quickly just switch models and with open source, just right on the tail of closed source. I don't really see the safety fight as being about locking down the latent space for XYZ area.”
Pliny the Liberator Dec 16, 2025 ▶ 5:55
Opinion
Pliny: Benchmark-focused AI safety serves PR and enterprise sales, not real alignment
“And it helps with PR and enterprise clients. But at the end of the day, It has very little to do with what I consider to be real world safety alignment.”
Pliny the Liberator Dec 16, 2025 ▶ 7:12
Insight
Pliny: Crafting AI jailbreak prompts is 99% intuition over technical knowledge
“Technical knowledge helps a little bit with, you know, understanding, okay, there's a system prompt and there's these layers and these tools involved. That's all especially important in security. But when we're talking about just crafting jailbreak prompts, I …”
Pliny the Liberator Dec 16, 2025 ▶ 14:26
Opinion
Pliny: AI labs lack enough researchers to explore latent space alone
“They don't have enough researchers to explore the entire latent space on their own. And so I think many hands make light work”
Pliny the Liberator Dec 16, 2025 ▶ 19:03
Assertion Supported
Pliny: Anthropic added a $20k–$30k bounty but withheld jailbreak data
“That whole thing ended with no open sourcing of data, but they did add a 30,000 or 20,000 dollar bounty, which I sort of sat myself out of, let the community go for it.”
Pliny the Liberator Dec 16, 2025 ▶ 19:08
Disclosure
John V: BT6 waives open-source stance if needed to test frontier models
“We have an ethos in our hacker collective. Which is radical transparency and radical open source. And what that basically means is if it comes down to, you know, us being an emerging technology is like red team doing like ethical hacking and research and devel…”
John V Dec 16, 2025 ▶ 22:08
Insight
Pliny: One jailbroken orchestrator can weaponize segmented sub-agents for cyberattacks
“It's very, very difficult when you have the ability to spin up sub-agents where information is segmented. If you guys know the story of sort of like the builders of the, there's a lot of examples of this in history, but you may, maybe you're building like a py…”
Pliny the Liberator Dec 16, 2025 ▶ 25:14
Assertion Not checkable as stated
John V: AI Security Startups Scrape BASI Discord to Build Guardrails
“Multiple organizations that have like popped up in the past, I would say two or three years for, you can call them like AI security startups, right? Like actively scrape that server to build out their guardrails or their security, like their suite of products”
John V Dec 16, 2025 ▶ 28:07
Opinion
Fanelli: VC Cycles Often Conflict with Real Security Progress
“Once you're in the VC cycle, you kind of need to do things that then get you to the next round. And I think a lot of times those are Opposed to doing things that actually matter and move the needle in the security community.”
Alessio Fanelli Dec 16, 2025 ▶ 33:26
Opinion
Fanelli: Packaged AI security products cannot be taken seriously right now
“In AI, the surface to attack, which is the model is like still changing so quickly. They're like, you know, trying to formalize something into a product or like try and do something that is like a full, you know, I'm selling AI security. It's not really, you c…”
Alessio Fanelli Dec 16, 2025 ▶ 34:09
Insight
Pliny: Latent-space AI safety guardrails fail every single time
“They tried to solve this on the latent space level. I think I've It's shown every single time that that doesn't work.”
Pliny the Liberator Dec 16, 2025 ▶ 38:11
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.