Cherny: Anthropic Can No Longer Demonstrate Prompt Injection on Opus 5
Boris Cherny · Boris Cherny: We Cut 80% of Claude Code’s Prompt · Y Combinator · Jul 27, 2026 · at 2:46
Boris Cherny, Head of Claude Code at Anthropic, explains Anthropic's mechanistic interpretability-based defenses against prompt injection.
“So essentially, if you combine a well-aligned model, so this is, like, essentially three years of research into alignment, With a prompt injection classifier, which we run for all traffic, and what this is doing is it's based on Chrysola's mechanistic interpretability work, where it's literally, we're looking at neurons in the model's brain that light up when prompt injection happens. So the model won't even tell you, but we can actually see those neurons, and we can figure out and diagnose that it's happening, and then you combine that with the auto mode classifier, and with these three layers, We just cannot demonstrate prompt injection anymore.”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →