Dec 16, 2025 · 40m · latent-space
⚡️Jailbreaking AGI: Pliny the Liberator & John V on Red Teaming, BT6, and the Future of AI Security
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
AI red teamers Pliny the Liberator and John V join the Latent Space podcast to discuss universal jailbreaking techniques, the mechanics of multi-agent adversarial attacks, and the imperative for cognitive freedom over corporate safety theater.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 22.8% of the talking time here. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Pliny aggressively dismisses industry attempts to tie latent space guardrails to safety, calling it a waste of time and lobotomization.
Hardest push from the hosts ▶ 15:43 Swyx Defends Academic ResearchersSwyx steps in to push back against John V's dismissive remarks regarding Anthropic researchers publishing multi-turn jailbreak papers.
Biggest teaching moment ▶ 36:22 John V Dismantles Single-Model Security ParadigmsJohn V educates the hosts on why securing AI cannot focus solely on text generation, explaining that tool and browser integrations create the true attack surface.
The host holds their own ▶ 33:22 Alessio Articulates Cybersecurity Investment RealitiesAlessio counters the guests' critique of VC incentives by citing real portfolio experience with tools like Metasploit and the unique constraints of cyber funding.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Universal Jailbreaks, Guardrails, and Security Theater | 6 | 5 | 4 | 2 | Alessio demonstrates domain knowledge by asking about refusal rate benchmarks and post-training interventions. Pliny forcefully reframes corporate safety guardrails as futile security theater and model lobotomization against an infinite Library of Babel search space. | |
| Anatomy of a Jailbreak: Libertas and Latent Space Seeds | 4 | 6 | 2 | 1 | The hosts bring up the Libertas repository, prompting Pliny to explain token stream disruption and latent space seeds. John V and Pliny educate the hosts on how mathematical quotient dividers steer multi-turn probability distributions. | |
| Soft Jailbreaks and the Anthropic Challenge Incident | 5 | 5 | 5 | 3 | John V mocks Anthropic for recently 'discovering' multi-turn jailbreaks, prompting Swyx to gently defend academic fellowships. Pliny details his public standoff with Anthropic over UI bugs and their refusal to open-source community research data. | |
| Autonomous Red Teaming and Sub-Agent Weaponization | 6 | 5 | 2 | 1 | Alessio steers the conversation to actual weaponization and AI-orchestrated attacks. Pliny breaks down how decomposing malicious workflows across isolated sub-agents bypasses individual agent safety filters. | |
| Cultivating the Hacker Collective: BASI, BT6, and Open Ecosystems | 7 | 3 | 4 | 4 | John V critiques Silicon Valley venture capital incentives for stifling radical open-source exploration. Alessio leverages his venture experience and portfolio precedents to explain why cybersecurity venture funding struggles with fast-moving model surfaces. | |
| Full-Stack AI Security versus Latent Space Safety | 5 | 7 | 4 | 1 | Pliny and John V explain that AI security must protect the entire system stack and meatspace execution layer rather than attempting to censor model weights. They argue that latent space safety fixes consistently fail compared to robust endpoint and tool-call boundary defenses. |