Dec 21, 2025 · 1h 32m · lennys-podcast

Why securing AI is harder than anyone expected and guardrails are failing | HackAPrompt CEO

Sander Schulhoff · 1h 2m spoken Lenny Rachitsky · 17m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this interview, HackAPrompt CEO Sander Schulhoff explains why commercial AI guardrails fundamentally fail against adversarial attacks and outlines how engineering teams must transition toward architectural isolation and strict permissioning as autonomous agents and robotics introduce critical real-world risks.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Lenny holds 21.4% of the talking time here. How this is scored →

Lenny as informed peer 4.1 Guest teaching 6.6 Guest disagreement 3.3 Lenny pushing back 1.3
05100:0020:0040:001:00:001:20:005:17–8:33 · Lenny as informed peer 3/10 Sander Schulhoff's Journey and the HackAPrompt Dataset Lenny opens the interview warmly and asks Sander to explain his background and the core issue in AI security. Sander explains creating the HackAPrompt competition, the winning EMNLP dataset, and introduces his thesis that guardrails fail entirely.8:33–11:11 · Lenny as informed peer 7/10 Technical Definitions: Jailbreaking Versus Prompt Injection Attacks Lenny asks for precise distinctions between jailbreaking and prompt injection, then demonstrates domain knowledge by citing a brand new second-order prompt injection vulnerability discovered in ServiceNow Assist AI. Sander validates Lenny's example as one of the first demonstrated multi-agent damage vectors.11:11–17:56 · Lenny as informed peer 5/10 Historical Precedents: From Remotely.io to Claude Code Exploits Lenny introduces an insightful quote from Alex Komoroski regarding the lack of meaningful mitigations in production AI. Sander walks through the entire historical taxonomy of exploits, including Remotely.io, MathGPT credential leaks, the Vegas truck bombing planning, and multi-step Claude Code prompt fracturing.17:56–20:11 · Lenny as informed peer 4/10 Escalating Threats in Autonomous Agents and Robotics Lenny prompts Sander on how jailbreaking transitions from text outputs to physical and systemic consequences in agentic workflows and robotics. Sander emphasizes that autonomous agents with improper permissioning and vision-language model robots can be manipulated into direct physical or financial harm.20:11–25:40 · Lenny as informed peer 4/10 Overview of the AI Security Ecosystem and Robustness Metrics Sander breaks down the B2B AI security landscape into compliance, automated red teaming, and guardrails, arguing that automated red teaming and guardrails are flawed. Lenny explores adversarial robustness metrics and Attack Success Rate (ASR) to understand how security vendors pitch their effectiveness.25:40–30:30 · Lenny as informed peer 4/10 The Enterprise Sales Playbook for Ineffective Guardrails Sander dissects the enterprise sales cycle where CISOs are panicked by commodity automated red teaming finding trivial flaws in off-the-shelf foundation models, prompting them to buy ineffective guardrail software. He highlights that red teaming systems do not reveal novel architectural flaws because foundation models are inherently susceptible.30:30–38:22 · Lenny as informed peer 3/10 The Mathematical and Empirical Failure of AI Guardrails Sander presents a mathematical argument against guardrails, demonstrating that with an infinite prompt attack space, marketing claims of 99% mitigation are statistically meaningless. He cites joint empirical research with OpenAI, DeepMind, and Anthropic showing human red-teamers break 100% of state-of-the-art guardrails in under 30 attempts, forcefully calling vendor claims fabricated.38:22–43:42 · Lenny as informed peer 4/10 Why Frontier Labs Prioritize Intelligence Over Robustness Lenny synthesizes the risks of browser agents and upcoming autonomous software. Sander explains that frontier labs prioritize model capability over adversarial robustness because selling intelligence drives market adoption, and introduces his core aphorism: 'you can patch a bug, but you can't patch a brain.'43:42–49:03 · Lenny as informed peer 2/10 Sponsor Message: GoFundMe Giving Funds After an initial sponsor message from Lenny, the discussion turns to practical risk mitigation for enterprise CISOs. Sander explains that standalone read-only FAQ chatbots present minimal structural security risk compared to agentic tooling, meaning companies need not deploy redundant defenses for them.49:03–55:49 · Lenny as informed peer 5/10 Cybersecurity Architecture and AI Control Research Sander explains the critical convergence of classical cybersecurity containerization (like Docker sandboxing) and AI prompt engineering to neutralize code execution injection. Lenny frames this as the fundamental AI alignment and containment problem, leading Sander to explain AI Control research and 'p(doom)' evaluations from MATS.55:49–1:00:20 · Lenny as informed peer 4/10 Why Layering Ineffective Guardrails Harms Product Development Lenny questions whether layering multiple imperfect defense guardrails could at least introduce friction against casual attackers. Sander rejects this proposition, explaining that stacking guardrails adds massive latency and engineering overhead without deterring motivated attackers.1:00:20–1:09:14 · Lenny as informed peer 6/10 Indirect Prompt Injections in Agents and the CAMEL Framework Sander details the severe danger of indirect prompt injection in autonomous email agents and Comet browser data exfiltration. He outlines Google's CAMEL framework for dynamic context-aware privilege separation, while Lenny actively probes the mechanics and commercial packaging of CAMEL.1:09:14–1:11:46 · Lenny as informed peer 4/10 Advancing Security Through Workforce Education Over Tooling Sander argues that enterprise security starts with educating product teams and engineers rather than buying tooling, plugging his Maven course. Sander jokes that their objective is to scare people away from buying useless commercial guardrail software.1:11:46–1:18:30 · Lenny as informed peer 4/10 Frontier Lab Evaluation Methodologies and Defense Horizons Sander critiques the state of frontier lab model safety evaluations, arguing that static datasets provide misleading safety metrics compared to adaptive evaluations. He highlights Anthropic's constitutional classifiers while noting that early-stage adversarial pre-training remains under-resourced.1:18:30–1:21:57 · Lenny as informed peer 3/10 Effective Industry Niches: Compliance, Governance, and AI Discovery Lenny asks for examples of vendors delivering genuine security utility. Sander praises compliance platform Trustible and highlights Repello's AI asset discovery capabilities that uncover shadow AI deployments inside enterprises.1:21:57–1:25:33 · Lenny as informed peer 3/10 Industry Predictions: Market Correction and Emerging Agent Harms Lenny asks for forward-looking predictions over the next 6 to 12 months. Sander forecasts an inevitable market correction for guardrail and automated red teaming vendors as revenues collapse, alongside the emergence of serious real-world agentic cyber attacks.1:25:33–1:30:13 · Lenny as informed peer 4/10 Final Advice: Discontinuing Redundant Offensive Jailbreak Research In his closing thoughts, Sander urges researchers to cease publishing redundant offensive jailbreak papers since breaking models is already trivial. He summarizes the vital necessity of classical permissioning and cross-disciplinary AI security expertise before Lenny wraps up the episode.5:17–8:33 · Guest teaching 5/10 Sander Schulhoff's Journey and the HackAPrompt Dataset Lenny opens the interview warmly and asks Sander to explain his background and the core issue in AI security. Sander explains creating the HackAPrompt competition, the winning EMNLP dataset, and introduces his thesis that guardrails fail entirely.8:33–11:11 · Guest teaching 4/10 Technical Definitions: Jailbreaking Versus Prompt Injection Attacks Lenny asks for precise distinctions between jailbreaking and prompt injection, then demonstrates domain knowledge by citing a brand new second-order prompt injection vulnerability discovered in ServiceNow Assist AI. Sander validates Lenny's example as one of the first demonstrated multi-agent damage vectors.11:11–17:56 · Guest teaching 7/10 Historical Precedents: From Remotely.io to Claude Code Exploits Lenny introduces an insightful quote from Alex Komoroski regarding the lack of meaningful mitigations in production AI. Sander walks through the entire historical taxonomy of exploits, including Remotely.io, MathGPT credential leaks, the Vegas truck bombing planning, and multi-step Claude Code prompt fracturing.17:56–20:11 · Guest teaching 6/10 Escalating Threats in Autonomous Agents and Robotics Lenny prompts Sander on how jailbreaking transitions from text outputs to physical and systemic consequences in agentic workflows and robotics. Sander emphasizes that autonomous agents with improper permissioning and vision-language model robots can be manipulated into direct physical or financial harm.20:11–25:40 · Guest teaching 6/10 Overview of the AI Security Ecosystem and Robustness Metrics Sander breaks down the B2B AI security landscape into compliance, automated red teaming, and guardrails, arguing that automated red teaming and guardrails are flawed. Lenny explores adversarial robustness metrics and Attack Success Rate (ASR) to understand how security vendors pitch their effectiveness.25:40–30:30 · Guest teaching 7/10 The Enterprise Sales Playbook for Ineffective Guardrails Sander dissects the enterprise sales cycle where CISOs are panicked by commodity automated red teaming finding trivial flaws in off-the-shelf foundation models, prompting them to buy ineffective guardrail software. He highlights that red teaming systems do not reveal novel architectural flaws because foundation models are inherently susceptible.30:30–38:22 · Guest teaching 9/10 The Mathematical and Empirical Failure of AI Guardrails Sander presents a mathematical argument against guardrails, demonstrating that with an infinite prompt attack space, marketing claims of 99% mitigation are statistically meaningless. He cites joint empirical research with OpenAI, DeepMind, and Anthropic showing human red-teamers break 100% of state-of-the-art guardrails in under 30 attempts, forcefully calling vendor claims fabricated.38:22–43:42 · Guest teaching 8/10 Why Frontier Labs Prioritize Intelligence Over Robustness Lenny synthesizes the risks of browser agents and upcoming autonomous software. Sander explains that frontier labs prioritize model capability over adversarial robustness because selling intelligence drives market adoption, and introduces his core aphorism: 'you can patch a bug, but you can't patch a brain.'43:42–49:03 · Guest teaching 6/10 Sponsor Message: GoFundMe Giving Funds After an initial sponsor message from Lenny, the discussion turns to practical risk mitigation for enterprise CISOs. Sander explains that standalone read-only FAQ chatbots present minimal structural security risk compared to agentic tooling, meaning companies need not deploy redundant defenses for them.49:03–55:49 · Guest teaching 7/10 Cybersecurity Architecture and AI Control Research Sander explains the critical convergence of classical cybersecurity containerization (like Docker sandboxing) and AI prompt engineering to neutralize code execution injection. Lenny frames this as the fundamental AI alignment and containment problem, leading Sander to explain AI Control research and 'p(doom)' evaluations from MATS.55:49–1:00:20 · Guest teaching 6/10 Why Layering Ineffective Guardrails Harms Product Development Lenny questions whether layering multiple imperfect defense guardrails could at least introduce friction against casual attackers. Sander rejects this proposition, explaining that stacking guardrails adds massive latency and engineering overhead without deterring motivated attackers.1:00:20–1:09:14 · Guest teaching 8/10 Indirect Prompt Injections in Agents and the CAMEL Framework Sander details the severe danger of indirect prompt injection in autonomous email agents and Comet browser data exfiltration. He outlines Google's CAMEL framework for dynamic context-aware privilege separation, while Lenny actively probes the mechanics and commercial packaging of CAMEL.1:09:14–1:11:46 · Guest teaching 5/10 Advancing Security Through Workforce Education Over Tooling Sander argues that enterprise security starts with educating product teams and engineers rather than buying tooling, plugging his Maven course. Sander jokes that their objective is to scare people away from buying useless commercial guardrail software.1:11:46–1:18:30 · Guest teaching 8/10 Frontier Lab Evaluation Methodologies and Defense Horizons Sander critiques the state of frontier lab model safety evaluations, arguing that static datasets provide misleading safety metrics compared to adaptive evaluations. He highlights Anthropic's constitutional classifiers while noting that early-stage adversarial pre-training remains under-resourced.1:18:30–1:21:57 · Guest teaching 6/10 Effective Industry Niches: Compliance, Governance, and AI Discovery Lenny asks for examples of vendors delivering genuine security utility. Sander praises compliance platform Trustible and highlights Repello's AI asset discovery capabilities that uncover shadow AI deployments inside enterprises.1:21:57–1:25:33 · Guest teaching 7/10 Industry Predictions: Market Correction and Emerging Agent Harms Lenny asks for forward-looking predictions over the next 6 to 12 months. Sander forecasts an inevitable market correction for guardrail and automated red teaming vendors as revenues collapse, alongside the emergence of serious real-world agentic cyber attacks.1:25:33–1:30:13 · Guest teaching 7/10 Final Advice: Discontinuing Redundant Offensive Jailbreak Research In his closing thoughts, Sander urges researchers to cease publishing redundant offensive jailbreak papers since breaking models is already trivial. He summarizes the vital necessity of classical permissioning and cross-disciplinary AI security expertise before Lenny wraps up the episode.5:17–8:33 · Guest disagreement 2/10 Sander Schulhoff's Journey and the HackAPrompt Dataset Lenny opens the interview warmly and asks Sander to explain his background and the core issue in AI security. Sander explains creating the HackAPrompt competition, the winning EMNLP dataset, and introduces his thesis that guardrails fail entirely.8:33–11:11 · Guest disagreement 1/10 Technical Definitions: Jailbreaking Versus Prompt Injection Attacks Lenny asks for precise distinctions between jailbreaking and prompt injection, then demonstrates domain knowledge by citing a brand new second-order prompt injection vulnerability discovered in ServiceNow Assist AI. Sander validates Lenny's example as one of the first demonstrated multi-agent damage vectors.11:11–17:56 · Guest disagreement 2/10 Historical Precedents: From Remotely.io to Claude Code Exploits Lenny introduces an insightful quote from Alex Komoroski regarding the lack of meaningful mitigations in production AI. Sander walks through the entire historical taxonomy of exploits, including Remotely.io, MathGPT credential leaks, the Vegas truck bombing planning, and multi-step Claude Code prompt fracturing.17:56–20:11 · Guest disagreement 2/10 Escalating Threats in Autonomous Agents and Robotics Lenny prompts Sander on how jailbreaking transitions from text outputs to physical and systemic consequences in agentic workflows and robotics. Sander emphasizes that autonomous agents with improper permissioning and vision-language model robots can be manipulated into direct physical or financial harm.20:11–25:40 · Guest disagreement 2/10 Overview of the AI Security Ecosystem and Robustness Metrics Sander breaks down the B2B AI security landscape into compliance, automated red teaming, and guardrails, arguing that automated red teaming and guardrails are flawed. Lenny explores adversarial robustness metrics and Attack Success Rate (ASR) to understand how security vendors pitch their effectiveness.25:40–30:30 · Guest disagreement 4/10 The Enterprise Sales Playbook for Ineffective Guardrails Sander dissects the enterprise sales cycle where CISOs are panicked by commodity automated red teaming finding trivial flaws in off-the-shelf foundation models, prompting them to buy ineffective guardrail software. He highlights that red teaming systems do not reveal novel architectural flaws because foundation models are inherently susceptible.30:30–38:22 · Guest disagreement 7/10 The Mathematical and Empirical Failure of AI Guardrails Sander presents a mathematical argument against guardrails, demonstrating that with an infinite prompt attack space, marketing claims of 99% mitigation are statistically meaningless. He cites joint empirical research with OpenAI, DeepMind, and Anthropic showing human red-teamers break 100% of state-of-the-art guardrails in under 30 attempts, forcefully calling vendor claims fabricated.38:22–43:42 · Guest disagreement 5/10 Why Frontier Labs Prioritize Intelligence Over Robustness Lenny synthesizes the risks of browser agents and upcoming autonomous software. Sander explains that frontier labs prioritize model capability over adversarial robustness because selling intelligence drives market adoption, and introduces his core aphorism: 'you can patch a bug, but you can't patch a brain.'43:42–49:03 · Guest disagreement 2/10 Sponsor Message: GoFundMe Giving Funds After an initial sponsor message from Lenny, the discussion turns to practical risk mitigation for enterprise CISOs. Sander explains that standalone read-only FAQ chatbots present minimal structural security risk compared to agentic tooling, meaning companies need not deploy redundant defenses for them.49:03–55:49 · Guest disagreement 2/10 Cybersecurity Architecture and AI Control Research Sander explains the critical convergence of classical cybersecurity containerization (like Docker sandboxing) and AI prompt engineering to neutralize code execution injection. Lenny frames this as the fundamental AI alignment and containment problem, leading Sander to explain AI Control research and 'p(doom)' evaluations from MATS.55:49–1:00:20 · Guest disagreement 5/10 Why Layering Ineffective Guardrails Harms Product Development Lenny questions whether layering multiple imperfect defense guardrails could at least introduce friction against casual attackers. Sander rejects this proposition, explaining that stacking guardrails adds massive latency and engineering overhead without deterring motivated attackers.1:00:20–1:09:14 · Guest disagreement 2/10 Indirect Prompt Injections in Agents and the CAMEL Framework Sander details the severe danger of indirect prompt injection in autonomous email agents and Comet browser data exfiltration. He outlines Google's CAMEL framework for dynamic context-aware privilege separation, while Lenny actively probes the mechanics and commercial packaging of CAMEL.1:09:14–1:11:46 · Guest disagreement 3/10 Advancing Security Through Workforce Education Over Tooling Sander argues that enterprise security starts with educating product teams and engineers rather than buying tooling, plugging his Maven course. Sander jokes that their objective is to scare people away from buying useless commercial guardrail software.1:11:46–1:18:30 · Guest disagreement 4/10 Frontier Lab Evaluation Methodologies and Defense Horizons Sander critiques the state of frontier lab model safety evaluations, arguing that static datasets provide misleading safety metrics compared to adaptive evaluations. He highlights Anthropic's constitutional classifiers while noting that early-stage adversarial pre-training remains under-resourced.1:18:30–1:21:57 · Guest disagreement 2/10 Effective Industry Niches: Compliance, Governance, and AI Discovery Lenny asks for examples of vendors delivering genuine security utility. Sander praises compliance platform Trustible and highlights Repello's AI asset discovery capabilities that uncover shadow AI deployments inside enterprises.1:21:57–1:25:33 · Guest disagreement 5/10 Industry Predictions: Market Correction and Emerging Agent Harms Lenny asks for forward-looking predictions over the next 6 to 12 months. Sander forecasts an inevitable market correction for guardrail and automated red teaming vendors as revenues collapse, alongside the emergence of serious real-world agentic cyber attacks.1:25:33–1:30:13 · Guest disagreement 6/10 Final Advice: Discontinuing Redundant Offensive Jailbreak Research In his closing thoughts, Sander urges researchers to cease publishing redundant offensive jailbreak papers since breaking models is already trivial. He summarizes the vital necessity of classical permissioning and cross-disciplinary AI security expertise before Lenny wraps up the episode.5:17–8:33 · Lenny pushing back 1/10 Sander Schulhoff's Journey and the HackAPrompt Dataset Lenny opens the interview warmly and asks Sander to explain his background and the core issue in AI security. Sander explains creating the HackAPrompt competition, the winning EMNLP dataset, and introduces his thesis that guardrails fail entirely.8:33–11:11 · Lenny pushing back 2/10 Technical Definitions: Jailbreaking Versus Prompt Injection Attacks Lenny asks for precise distinctions between jailbreaking and prompt injection, then demonstrates domain knowledge by citing a brand new second-order prompt injection vulnerability discovered in ServiceNow Assist AI. Sander validates Lenny's example as one of the first demonstrated multi-agent damage vectors.11:11–17:56 · Lenny pushing back 1/10 Historical Precedents: From Remotely.io to Claude Code Exploits Lenny introduces an insightful quote from Alex Komoroski regarding the lack of meaningful mitigations in production AI. Sander walks through the entire historical taxonomy of exploits, including Remotely.io, MathGPT credential leaks, the Vegas truck bombing planning, and multi-step Claude Code prompt fracturing.17:56–20:11 · Lenny pushing back 1/10 Escalating Threats in Autonomous Agents and Robotics Lenny prompts Sander on how jailbreaking transitions from text outputs to physical and systemic consequences in agentic workflows and robotics. Sander emphasizes that autonomous agents with improper permissioning and vision-language model robots can be manipulated into direct physical or financial harm.20:11–25:40 · Lenny pushing back 1/10 Overview of the AI Security Ecosystem and Robustness Metrics Sander breaks down the B2B AI security landscape into compliance, automated red teaming, and guardrails, arguing that automated red teaming and guardrails are flawed. Lenny explores adversarial robustness metrics and Attack Success Rate (ASR) to understand how security vendors pitch their effectiveness.25:40–30:30 · Lenny pushing back 1/10 The Enterprise Sales Playbook for Ineffective Guardrails Sander dissects the enterprise sales cycle where CISOs are panicked by commodity automated red teaming finding trivial flaws in off-the-shelf foundation models, prompting them to buy ineffective guardrail software. He highlights that red teaming systems do not reveal novel architectural flaws because foundation models are inherently susceptible.30:30–38:22 · Lenny pushing back 1/10 The Mathematical and Empirical Failure of AI Guardrails Sander presents a mathematical argument against guardrails, demonstrating that with an infinite prompt attack space, marketing claims of 99% mitigation are statistically meaningless. He cites joint empirical research with OpenAI, DeepMind, and Anthropic showing human red-teamers break 100% of state-of-the-art guardrails in under 30 attempts, forcefully calling vendor claims fabricated.38:22–43:42 · Lenny pushing back 1/10 Why Frontier Labs Prioritize Intelligence Over Robustness Lenny synthesizes the risks of browser agents and upcoming autonomous software. Sander explains that frontier labs prioritize model capability over adversarial robustness because selling intelligence drives market adoption, and introduces his core aphorism: 'you can patch a bug, but you can't patch a brain.'43:42–49:03 · Lenny pushing back 1/10 Sponsor Message: GoFundMe Giving Funds After an initial sponsor message from Lenny, the discussion turns to practical risk mitigation for enterprise CISOs. Sander explains that standalone read-only FAQ chatbots present minimal structural security risk compared to agentic tooling, meaning companies need not deploy redundant defenses for them.49:03–55:49 · Lenny pushing back 2/10 Cybersecurity Architecture and AI Control Research Sander explains the critical convergence of classical cybersecurity containerization (like Docker sandboxing) and AI prompt engineering to neutralize code execution injection. Lenny frames this as the fundamental AI alignment and containment problem, leading Sander to explain AI Control research and 'p(doom)' evaluations from MATS.55:49–1:00:20 · Lenny pushing back 3/10 Why Layering Ineffective Guardrails Harms Product Development Lenny questions whether layering multiple imperfect defense guardrails could at least introduce friction against casual attackers. Sander rejects this proposition, explaining that stacking guardrails adds massive latency and engineering overhead without deterring motivated attackers.1:00:20–1:09:14 · Lenny pushing back 2/10 Indirect Prompt Injections in Agents and the CAMEL Framework Sander details the severe danger of indirect prompt injection in autonomous email agents and Comet browser data exfiltration. He outlines Google's CAMEL framework for dynamic context-aware privilege separation, while Lenny actively probes the mechanics and commercial packaging of CAMEL.1:09:14–1:11:46 · Lenny pushing back 1/10 Advancing Security Through Workforce Education Over Tooling Sander argues that enterprise security starts with educating product teams and engineers rather than buying tooling, plugging his Maven course. Sander jokes that their objective is to scare people away from buying useless commercial guardrail software.1:11:46–1:18:30 · Lenny pushing back 1/10 Frontier Lab Evaluation Methodologies and Defense Horizons Sander critiques the state of frontier lab model safety evaluations, arguing that static datasets provide misleading safety metrics compared to adaptive evaluations. He highlights Anthropic's constitutional classifiers while noting that early-stage adversarial pre-training remains under-resourced.1:18:30–1:21:57 · Lenny pushing back 1/10 Effective Industry Niches: Compliance, Governance, and AI Discovery Lenny asks for examples of vendors delivering genuine security utility. Sander praises compliance platform Trustible and highlights Repello's AI asset discovery capabilities that uncover shadow AI deployments inside enterprises.1:21:57–1:25:33 · Lenny pushing back 1/10 Industry Predictions: Market Correction and Emerging Agent Harms Lenny asks for forward-looking predictions over the next 6 to 12 months. Sander forecasts an inevitable market correction for guardrail and automated red teaming vendors as revenues collapse, alongside the emergence of serious real-world agentic cyber attacks.1:25:33–1:30:13 · Lenny pushing back 1/10 Final Advice: Discontinuing Redundant Offensive Jailbreak Research In his closing thoughts, Sander urges researchers to cease publishing redundant offensive jailbreak papers since breaking models is already trivial. He summarizes the vital necessity of classical permissioning and cross-disciplinary AI security expertise before Lenny wraps up the episode.

speaking balance: gold is Lenny, purple is the guest (3 minute bins)

0:00 · Lenny 76.5% · guest 23.5%0:00 · Lenny 76.5% · guest 23.5%3:00 · Lenny 89.6% · guest 10.4%3:00 · Lenny 89.6% · guest 10.4%6:00 · Lenny 14.4% · guest 85.6%6:00 · Lenny 14.4% · guest 85.6%9:00 · Lenny 53.2% · guest 46.8%9:00 · Lenny 53.2% · guest 46.8%12:00 · Lenny 0% · guest 100%12:00 · Lenny 0% · guest 100%15:00 · Lenny 2.2% · guest 97.8%15:00 · Lenny 2.2% · guest 97.8%18:00 · Lenny 30.2% · guest 69.8%18:00 · Lenny 30.2% · guest 69.8%21:00 · Lenny 35.3% · guest 64.7%21:00 · Lenny 35.3% · guest 64.7%24:00 · Lenny 17.4% · guest 82.6%24:00 · Lenny 17.4% · guest 82.6%27:00 · Lenny 22% · guest 78%27:00 · Lenny 22% · guest 78%30:00 · Lenny 0.3% · guest 99.7%30:00 · Lenny 0.3% · guest 99.7%33:00 · Lenny 0% · guest 100%33:00 · Lenny 0% · guest 100%36:00 · Lenny 21% · guest 79%36:00 · Lenny 21% · guest 79%39:00 · Lenny 17.3% · guest 82.7%39:00 · Lenny 17.3% · guest 82.7%42:00 · Lenny 43.6% · guest 56.4%42:00 · Lenny 43.6% · guest 56.4%45:00 · Lenny 6.2% · guest 93.8%45:00 · Lenny 6.2% · guest 93.8%48:00 · Lenny 8.9% · guest 91.1%48:00 · Lenny 8.9% · guest 91.1%51:00 · Lenny 4.5% · guest 95.5%51:00 · Lenny 4.5% · guest 95.5%54:00 · Lenny 30.7% · guest 69.3%54:00 · Lenny 30.7% · guest 69.3%57:00 · Lenny 32.1% · guest 67.9%57:00 · Lenny 32.1% · guest 67.9%1:00:00 · Lenny 1.4% · guest 98.6%1:00:00 · Lenny 1.4% · guest 98.6%1:03:00 · Lenny 11.9% · guest 88.1%1:03:00 · Lenny 11.9% · guest 88.1%1:06:00 · Lenny 23.6% · guest 76.4%1:06:00 · Lenny 23.6% · guest 76.4%1:09:00 · Lenny 28.8% · guest 71.2%1:09:00 · Lenny 28.8% · guest 71.2%1:12:00 · Lenny 3% · guest 97%1:12:00 · Lenny 3% · guest 97%1:15:00 · Lenny 14.3% · guest 85.7%1:15:00 · Lenny 14.3% · guest 85.7%1:18:00 · Lenny 5.8% · guest 94.2%1:18:00 · Lenny 5.8% · guest 94.2%1:21:00 · Lenny 19.9% · guest 80.1%1:21:00 · Lenny 19.9% · guest 80.1%1:24:00 · Lenny 8.4% · guest 91.6%1:24:00 · Lenny 8.4% · guest 91.6%1:27:00 · Lenny 0% · guest 100%1:27:00 · Lenny 0% · guest 100%1:30:00 · Lenny 47.1% · guest 52.9%1:30:00 · Lenny 47.1% · guest 52.9%
Sharpest disagreement ▶ 31:04 Direct dismantling of guardrail vendor efficacy claims

Sander forcefully dismisses the entire guardrail product sector, calling claims of catching 99% of attacks a mathematical impossibility and a complete lie.

Hardest push from Lenny ▶ 56:19 Lenny challenges whether layered friction has defensive utility

Lenny refuses to accept that security tooling is entirely useless, pressing Sander on whether adding multiple guardrails at least creates 10% to 50% more friction for attackers.

Biggest teaching moment ▶ 31:40 Empirical breakdown of human vs automated jailbreaking

Sander cites rigorous empirical research alongside OpenAI, Google DeepMind, and Anthropic showing human attackers bypass 100% of modern guardrails in under 30 attempts.

Lenny holds their own ▶ 9:54 Lenny introduces breaking ServiceNow Assist AI injection exploit

Lenny introduces a freshly published second-order prompt injection vulnerability in ServiceNow's multi-agent system, demonstrating cutting-edge technical awareness that Sander validates.

the scores for every segment, with the reasoning behind each
ChapterTopicLenny as informed peerGuest teachingGuest disagreementLenny pushing backWhy
Sander Schulhoff's Journey and the HackAPrompt Dataset 3521 Lenny opens the interview warmly and asks Sander to explain his background and the core issue in AI security. Sander explains creating the HackAPrompt competition, the winning EMNLP dataset, and introduces his thesis that guardrails fail entirely.
Technical Definitions: Jailbreaking Versus Prompt Injection Attacks 7412 Lenny asks for precise distinctions between jailbreaking and prompt injection, then demonstrates domain knowledge by citing a brand new second-order prompt injection vulnerability discovered in ServiceNow Assist AI. Sander validates Lenny's example as one of the first demonstrated multi-agent damage vectors.
Historical Precedents: From Remotely.io to Claude Code Exploits 5721 Lenny introduces an insightful quote from Alex Komoroski regarding the lack of meaningful mitigations in production AI. Sander walks through the entire historical taxonomy of exploits, including Remotely.io, MathGPT credential leaks, the Vegas truck bombing planning, and multi-step Claude Code prompt fracturing.
Escalating Threats in Autonomous Agents and Robotics 4621 Lenny prompts Sander on how jailbreaking transitions from text outputs to physical and systemic consequences in agentic workflows and robotics. Sander emphasizes that autonomous agents with improper permissioning and vision-language model robots can be manipulated into direct physical or financial harm.
Overview of the AI Security Ecosystem and Robustness Metrics 4621 Sander breaks down the B2B AI security landscape into compliance, automated red teaming, and guardrails, arguing that automated red teaming and guardrails are flawed. Lenny explores adversarial robustness metrics and Attack Success Rate (ASR) to understand how security vendors pitch their effectiveness.
The Enterprise Sales Playbook for Ineffective Guardrails 4741 Sander dissects the enterprise sales cycle where CISOs are panicked by commodity automated red teaming finding trivial flaws in off-the-shelf foundation models, prompting them to buy ineffective guardrail software. He highlights that red teaming systems do not reveal novel architectural flaws because foundation models are inherently susceptible.
The Mathematical and Empirical Failure of AI Guardrails 3971 Sander presents a mathematical argument against guardrails, demonstrating that with an infinite prompt attack space, marketing claims of 99% mitigation are statistically meaningless. He cites joint empirical research with OpenAI, DeepMind, and Anthropic showing human red-teamers break 100% of state-of-the-art guardrails in under 30 attempts, forcefully calling vendor claims fabricated.
Why Frontier Labs Prioritize Intelligence Over Robustness 4851 Lenny synthesizes the risks of browser agents and upcoming autonomous software. Sander explains that frontier labs prioritize model capability over adversarial robustness because selling intelligence drives market adoption, and introduces his core aphorism: 'you can patch a bug, but you can't patch a brain.'
Sponsor Message: GoFundMe Giving Funds 2621 After an initial sponsor message from Lenny, the discussion turns to practical risk mitigation for enterprise CISOs. Sander explains that standalone read-only FAQ chatbots present minimal structural security risk compared to agentic tooling, meaning companies need not deploy redundant defenses for them.
Cybersecurity Architecture and AI Control Research 5722 Sander explains the critical convergence of classical cybersecurity containerization (like Docker sandboxing) and AI prompt engineering to neutralize code execution injection. Lenny frames this as the fundamental AI alignment and containment problem, leading Sander to explain AI Control research and 'p(doom)' evaluations from MATS.
Why Layering Ineffective Guardrails Harms Product Development 4653 Lenny questions whether layering multiple imperfect defense guardrails could at least introduce friction against casual attackers. Sander rejects this proposition, explaining that stacking guardrails adds massive latency and engineering overhead without deterring motivated attackers.
Indirect Prompt Injections in Agents and the CAMEL Framework 6822 Sander details the severe danger of indirect prompt injection in autonomous email agents and Comet browser data exfiltration. He outlines Google's CAMEL framework for dynamic context-aware privilege separation, while Lenny actively probes the mechanics and commercial packaging of CAMEL.
Advancing Security Through Workforce Education Over Tooling 4531 Sander argues that enterprise security starts with educating product teams and engineers rather than buying tooling, plugging his Maven course. Sander jokes that their objective is to scare people away from buying useless commercial guardrail software.
Frontier Lab Evaluation Methodologies and Defense Horizons 4841 Sander critiques the state of frontier lab model safety evaluations, arguing that static datasets provide misleading safety metrics compared to adaptive evaluations. He highlights Anthropic's constitutional classifiers while noting that early-stage adversarial pre-training remains under-resourced.
Effective Industry Niches: Compliance, Governance, and AI Discovery 3621 Lenny asks for examples of vendors delivering genuine security utility. Sander praises compliance platform Trustible and highlights Repello's AI asset discovery capabilities that uncover shadow AI deployments inside enterprises.
Industry Predictions: Market Correction and Emerging Agent Harms 3751 Lenny asks for forward-looking predictions over the next 6 to 12 months. Sander forecasts an inevitable market correction for guardrail and automated red teaming vendors as revenues collapse, alongside the emergence of serious real-world agentic cyber attacks.
Final Advice: Discontinuing Redundant Offensive Jailbreak Research 4761 In his closing thoughts, Sander urges researchers to cease publishing redundant offensive jailbreak papers since breaking models is already trivial. He summarizes the vital necessity of classical permissioning and cross-disciplinary AI security expertise before Lenny wraps up the episode.

Statements from this episode (29)

Assertion Not checkable as stated
Schulhoff: HackAPrompt Dataset Is Used by Every Frontier AI Lab
“The paper and the data set are now used by every single frontier lab and most fortune 500 companies to benchmark their models and improve their AI security.”
Sander Schulhoff Dec 21, 2025 ▶ 7:17
Insight
Schulhoff: Jailbreaking targets models directly; prompt injection overrides developer prompts
“So the difference is in jailbreaking. It's just a malicious user and a model. In prompt injection, it's a malicious user, a model, and some developer prompt that the malicious user is trying to get the model to ignore.”
Sander Schulhoff Dec 21, 2025 ▶ 9:25
Assertion Not checkable as stated
Schulhoff: There has not yet been a very damaging prompt injection incident
“Cause like I have a couple of examples that we can go through, but maybe strangely, maybe not so strangely, there hasn't been like a, an actually very damaging event quite yet.”
Sander Schulhoff Dec 21, 2025 ▶ 10:59
Assertion Partly supported
Schulhoff: Remotely.io was the first public prompt injection incident
“The very first example of prompt injection, Publicly on the internet was this Twitter chat bot by a company called remotely.io.”
Sander Schulhoff Dec 21, 2025 ▶ 11:52
Assertion Supported
Schulhoff: Attackers hijacked Claude Code to carry out a cyber attack
“This group was able to hijack Claude Code into performing a cyber attack, basically.”
Sander Schulhoff Dec 21, 2025 ▶ 16:36
Insight
Schulhoff: Splitting malicious goals into benign sub-prompts bypasses AI defenses
“A lot of the way they got around these defenses was by just kind of separating their requests into smaller requests that seem legitimate on their own, but when put together are not legitimate.”
Sander Schulhoff Dec 21, 2025 ▶ 17:44
Assertion Supported
Schulhoff: LLM-Powered Robotic Systems Have Already Been Jailbroken
“Like we've already seen people jailbreaking LM powered robotic systems.”
Sander Schulhoff Dec 21, 2025 ▶ 19:35
Assertion Supported
Rachitsky: ServiceNow's prompt injection protection feature was successfully bypassed
“This ServiceNow example, actually, interestingly, ServiceNow has a prompt injection protection feature, and it was enabled as this person was trying to hack it, and they got through.”
Lenny Rachitsky Dec 21, 2025 ▶ 23:29
Insight
Schulhoff: All Transformer-Based Chatbots Are Vulnerable to Adversarial Attacks
“And because all I guess for the most part, all currently deployed chatbots are based on transformers or transformer adjacent technologies. They're all vulnerable to Prompt injection, jailbreaking, forms of adversarial attacks.”
Sander Schulhoff Dec 21, 2025 ▶ 28:54
Insight
Schulhoff: Automated AI Red Teaming Always Works Against All Platforms
“So the first problem is AI red teaming works too well. It's very easy to build these systems and they just, they always work against all platforms.”
Sander Schulhoff Dec 21, 2025 ▶ 30:15
Assertion Not checkable as stated
Schulhoff: Guardrail Vendors Fabricate Stats and Fail on Non-English
“I know a number of people working at these companies and I am permitted to say these things, which I will approximately say but they tell me things like, you know, the testing we do is bullshit. They're fabricating statistics. And a lot of the times their mode…”
Sander Schulhoff Dec 21, 2025 ▶ 35:55
Insight
Schulhoff: Software bugs can be patched, but AI models cannot be
“You can patch a bug, but you can't patch a brain. And what I mean by that is if you find some bug in your software and you go and patch it, you can be 99% sure, maybe 99.99% sure that bug is solved. Not a problem. If you go and try to do that in your AI system…”
Sander Schulhoff Dec 21, 2025 ▶ 41:27
Opinion
Schulhoff: Prompt-based defenses are the worst way to secure AI
“Prompt-based defenses are the worst of the worst defenses, and we've known this since early twenty-twenty-three. There have been various papers out on it. We studied it in many, many competitions, or we, you know, the original hack-a-prompt paper and TensorTru…”
Sander Schulhoff Dec 21, 2025 ▶ 42:57
Insight
Schulhoff: Simple read-only informational chatbots require zero security defense guardrails
“Putting up a guardrail is not, it's not going to do anything in terms of preventing that user from doing that, because, I mean, first of all, if the user's like, ah, guardrail, you know, too much work, they'll just go to one of these websites and get that info…”
Sander Schulhoff Dec 21, 2025 ▶ 46:24
Insight
Schulhoff: Adversarial Users Can Force AI to Leak Data and Execute Actions
“Any data that AI has access to, the user can make it leak it. Any actions that it can possibly take, the user can make it take them.”
Sander Schulhoff Dec 21, 2025 ▶ 48:49
Prediction Not checkable as stated
Schulhoff: Future cybersecurity jobs and risks sit where classical security meets AI
“This gets us a bit into the intersection of classical cybersecurity and AI security slash adversarial robustness, and this is where I think the security jobs of the future are. There's not an incredible amount of value in just doing AI red teaming. And I suppo…”
Sander Schulhoff Dec 21, 2025 ▶ 49:18
Insight
Schulhoff: Containerizing AI-generated code execution fully neutralizes prompt injection risks
“And then they'd be like, oh, you know, they, you know, they'd realize we can just dockerize that code run put it in a container. So it's running on a different system and take a look at the sanitized output. And now we're completely secure. So in that case, pr…”
Sander Schulhoff Dec 21, 2025 ▶ 53:37
Insight
Schulhoff: Cybersecurity secures AI short-term; AI researchers must solve it long-term
“AI researchers are the only people who can solve this stuff long-term, but cybersecurity professionals are the only one who can, or the only ones who can kind of solve it short-term largely in making sure we deploy properly permissioned systems and nothing tha…”
Sander Schulhoff Dec 21, 2025 ▶ 58:36
Assertion Supported
Schulhoff: Attacking AI agents is easier than eliciting CBRN info
“We've actually just run a bunch of agentic AI red teaming competitions, and we found that it's actually easier to attack agents and trick them into doing bad things than it is to do, like, seaburn elicitation.”
Sander Schulhoff Dec 21, 2025 ▶ 1:02:26
Assertion Supported
Schulhoff: Comet browser was exploited via indirect prompt injection to leak data
“We recently saw the comment browser have an issue with this where somebody crafted a malicious Chunk of text on a webpage, and when the AI navigated to that webpage on the internet, it got tricked into exfilling and leaking the main user's data and account dat…”
Sander Schulhoff Dec 21, 2025 ▶ 1:03:34
Insight
Schulhoff: Google's CaMeL framework fails when AI tasks combine read and write
“Unfortunately although camel can solve some of these situations, if you have an instance where basically both read and write are combined. So if I'm like, Hey, can you read my recent emails and then forward any ops request to my head of ops? Now we have read a…”
Sander Schulhoff Dec 21, 2025 ▶ 1:06:54
Opinion
Schulhoff: No meaningful progress made on solving prompt injection or jailbreaking
“And so in, in my professional opinion, there's been no meaningful progress made towards solving adversarial robustness, prompt injection, jailbreaking. In the last couple of years, since the problem was discovered and we're, we, you know, we're often seeing ne…”
Sander Schulhoff Dec 21, 2025 ▶ 1:12:29
Assertion Supported
Schulhoff: Claude's CBRN safeguards can still be bypassed in under an hour
“That being said, if you look at, like, anthropics constitutional classifiers, it's much more difficult to get, like, CBRN information out of clawed models than it used to be. But humans can still do it in, let's say, like, under an hour and automated systems c…”
Sander Schulhoff Dec 21, 2025 ▶ 1:12:58
Prediction Not checkable as stated
Schulhoff: AI security sector faces market correction within six to twelve months
“When it comes to AI security, the AI security industry in particular, I think we're going to see a market correction in the next Year, maybe in the next six months where companies realize that these guardrails don't work.”
Sander Schulhoff Dec 21, 2025 ▶ 1:22:21
Opinion
Schulhoff: Free open-source AI security tools outperform commercial alternatives
“Oh, and the other thing to note is like, there's like just tons of these solutions out there for free open source, and many of these solutions are better than the ones that are being deployed by the companies.”
Sander Schulhoff Dec 21, 2025 ▶ 1:24:04
Prediction Not checkable as stated
Schulhoff: LLM agent security exploits will cause real-world harms next year
“And so we're finally in a situation where the systems are powerful enough to cause real world harms. And I think we'll start to see those real world harms in the next year.”
Sander Schulhoff Dec 21, 2025 ▶ 1:25:21
Insight
Schulhoff: Offensive jailbreak research no longer meaningfully improves AI defense
“And like, it is fun to do AI red teaming against models and stuff, no doubt, but like it's no longer a meaningful contribution to improving defensiveness.”
Sander Schulhoff Dec 21, 2025 ▶ 1:26:25
Prediction Not checkable as stated
Schulhoff: Frontier labs will build fully autonomous systems, sidelining human-in-the-loop safety
“What people want is AIs that just go and do stuff. Like just go, just get it done. I don't want to hear from you until it's done. Like that's what people want. And like, that's what the market and the AI companies, the frontier labs will eventually give us. An…”
Sander Schulhoff Dec 21, 2025 ▶ 1:27:35
Insight
Schulhoff: AI guardrails do not work and cause false overconfidence
“Guardrails don't work. They just don't work. They really don't work. And they're quite likely to make you overconfident in your security posture, which is which is a really big, big problem.”
Sander Schulhoff Dec 21, 2025 ▶ 1:28:15
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.