Dec 21, 2025 · 1h 32m · lennys-podcast
Why securing AI is harder than anyone expected and guardrails are failing | HackAPrompt CEO
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this interview, HackAPrompt CEO Sander Schulhoff explains why commercial AI guardrails fundamentally fail against adversarial attacks and outlines how engineering teams must transition toward architectural isolation and strict permissioning as autonomous agents and robotics introduce critical real-world risks.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Lenny holds 21.4% of the talking time here. How this is scored →
speaking balance: gold is Lenny, purple is the guest (3 minute bins)
Sander forcefully dismisses the entire guardrail product sector, calling claims of catching 99% of attacks a mathematical impossibility and a complete lie.
Hardest push from Lenny ▶ 56:19 Lenny challenges whether layered friction has defensive utilityLenny refuses to accept that security tooling is entirely useless, pressing Sander on whether adding multiple guardrails at least creates 10% to 50% more friction for attackers.
Biggest teaching moment ▶ 31:40 Empirical breakdown of human vs automated jailbreakingSander cites rigorous empirical research alongside OpenAI, Google DeepMind, and Anthropic showing human attackers bypass 100% of modern guardrails in under 30 attempts.
Lenny holds their own ▶ 9:54 Lenny introduces breaking ServiceNow Assist AI injection exploitLenny introduces a freshly published second-order prompt injection vulnerability in ServiceNow's multi-agent system, demonstrating cutting-edge technical awareness that Sander validates.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Lenny as informed peer | Guest teaching | Guest disagreement | Lenny pushing back | Why |
|---|---|---|---|---|---|---|
| Sander Schulhoff's Journey and the HackAPrompt Dataset | 3 | 5 | 2 | 1 | Lenny opens the interview warmly and asks Sander to explain his background and the core issue in AI security. Sander explains creating the HackAPrompt competition, the winning EMNLP dataset, and introduces his thesis that guardrails fail entirely. | |
| Technical Definitions: Jailbreaking Versus Prompt Injection Attacks | 7 | 4 | 1 | 2 | Lenny asks for precise distinctions between jailbreaking and prompt injection, then demonstrates domain knowledge by citing a brand new second-order prompt injection vulnerability discovered in ServiceNow Assist AI. Sander validates Lenny's example as one of the first demonstrated multi-agent damage vectors. | |
| Historical Precedents: From Remotely.io to Claude Code Exploits | 5 | 7 | 2 | 1 | Lenny introduces an insightful quote from Alex Komoroski regarding the lack of meaningful mitigations in production AI. Sander walks through the entire historical taxonomy of exploits, including Remotely.io, MathGPT credential leaks, the Vegas truck bombing planning, and multi-step Claude Code prompt fracturing. | |
| Escalating Threats in Autonomous Agents and Robotics | 4 | 6 | 2 | 1 | Lenny prompts Sander on how jailbreaking transitions from text outputs to physical and systemic consequences in agentic workflows and robotics. Sander emphasizes that autonomous agents with improper permissioning and vision-language model robots can be manipulated into direct physical or financial harm. | |
| Overview of the AI Security Ecosystem and Robustness Metrics | 4 | 6 | 2 | 1 | Sander breaks down the B2B AI security landscape into compliance, automated red teaming, and guardrails, arguing that automated red teaming and guardrails are flawed. Lenny explores adversarial robustness metrics and Attack Success Rate (ASR) to understand how security vendors pitch their effectiveness. | |
| The Enterprise Sales Playbook for Ineffective Guardrails | 4 | 7 | 4 | 1 | Sander dissects the enterprise sales cycle where CISOs are panicked by commodity automated red teaming finding trivial flaws in off-the-shelf foundation models, prompting them to buy ineffective guardrail software. He highlights that red teaming systems do not reveal novel architectural flaws because foundation models are inherently susceptible. | |
| The Mathematical and Empirical Failure of AI Guardrails | 3 | 9 | 7 | 1 | Sander presents a mathematical argument against guardrails, demonstrating that with an infinite prompt attack space, marketing claims of 99% mitigation are statistically meaningless. He cites joint empirical research with OpenAI, DeepMind, and Anthropic showing human red-teamers break 100% of state-of-the-art guardrails in under 30 attempts, forcefully calling vendor claims fabricated. | |
| Why Frontier Labs Prioritize Intelligence Over Robustness | 4 | 8 | 5 | 1 | Lenny synthesizes the risks of browser agents and upcoming autonomous software. Sander explains that frontier labs prioritize model capability over adversarial robustness because selling intelligence drives market adoption, and introduces his core aphorism: 'you can patch a bug, but you can't patch a brain.' | |
| Sponsor Message: GoFundMe Giving Funds | 2 | 6 | 2 | 1 | After an initial sponsor message from Lenny, the discussion turns to practical risk mitigation for enterprise CISOs. Sander explains that standalone read-only FAQ chatbots present minimal structural security risk compared to agentic tooling, meaning companies need not deploy redundant defenses for them. | |
| Cybersecurity Architecture and AI Control Research | 5 | 7 | 2 | 2 | Sander explains the critical convergence of classical cybersecurity containerization (like Docker sandboxing) and AI prompt engineering to neutralize code execution injection. Lenny frames this as the fundamental AI alignment and containment problem, leading Sander to explain AI Control research and 'p(doom)' evaluations from MATS. | |
| Why Layering Ineffective Guardrails Harms Product Development | 4 | 6 | 5 | 3 | Lenny questions whether layering multiple imperfect defense guardrails could at least introduce friction against casual attackers. Sander rejects this proposition, explaining that stacking guardrails adds massive latency and engineering overhead without deterring motivated attackers. | |
| Indirect Prompt Injections in Agents and the CAMEL Framework | 6 | 8 | 2 | 2 | Sander details the severe danger of indirect prompt injection in autonomous email agents and Comet browser data exfiltration. He outlines Google's CAMEL framework for dynamic context-aware privilege separation, while Lenny actively probes the mechanics and commercial packaging of CAMEL. | |
| Advancing Security Through Workforce Education Over Tooling | 4 | 5 | 3 | 1 | Sander argues that enterprise security starts with educating product teams and engineers rather than buying tooling, plugging his Maven course. Sander jokes that their objective is to scare people away from buying useless commercial guardrail software. | |
| Frontier Lab Evaluation Methodologies and Defense Horizons | 4 | 8 | 4 | 1 | Sander critiques the state of frontier lab model safety evaluations, arguing that static datasets provide misleading safety metrics compared to adaptive evaluations. He highlights Anthropic's constitutional classifiers while noting that early-stage adversarial pre-training remains under-resourced. | |
| Effective Industry Niches: Compliance, Governance, and AI Discovery | 3 | 6 | 2 | 1 | Lenny asks for examples of vendors delivering genuine security utility. Sander praises compliance platform Trustible and highlights Repello's AI asset discovery capabilities that uncover shadow AI deployments inside enterprises. | |
| Industry Predictions: Market Correction and Emerging Agent Harms | 3 | 7 | 5 | 1 | Lenny asks for forward-looking predictions over the next 6 to 12 months. Sander forecasts an inevitable market correction for guardrail and automated red teaming vendors as revenues collapse, alongside the emergence of serious real-world agentic cyber attacks. | |
| Final Advice: Discontinuing Redundant Offensive Jailbreak Research | 4 | 7 | 6 | 1 | In his closing thoughts, Sander urges researchers to cease publishing redundant offensive jailbreak papers since breaking models is already trivial. He summarizes the vital necessity of classical permissioning and cross-disciplinary AI security expertise before Lenny wraps up the episode. |