Jun 19, 2025 · 1h 37m · lennys-podcast
AI prompt engineering in 2025: What works and what doesn’t | Sander Schulhoff
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this comprehensive interview, AI researcher Sander Schulhoff details proven prompt engineering techniques for production applications and examines the persistent security vulnerabilities and red-teaming challenges confronting autonomous AI systems.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Lenny holds 25.2% of the talking time here. How this is scored →
speaking balance: gold is Lenny, purple is the guest (3 minute bins)
Sander emphatically rejects the widely marketed industry consensus that guardrails or prompt-level instructions can secure LLMs, dismissing these commercial solutions as completely ineffective.
Hardest push from Lenny ▶ 1:15:07 Challenging whether prompt injection is solvableLenny directly challenges Sander on whether prompt injection is an engineering problem that can eventually be solved or an permanent arms race.
Biggest teaching moment ▶ 1:11:05 The intelligence gap in AI guardrailsSander explains the intelligence gap vulnerability, showing how encoded or obfuscated prompts easily bypass smaller guardrail models while tricking the main model.
Lenny holds their own ▶ 1:31:44 Angel investment reveal in Daylight ComputerLenny demonstrates his industry insider standing by revealing to Sander that he was an original angel investor in the hardware company Sander brought for show-and-tell.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Lenny as informed peer | Guest teaching | Guest disagreement | Lenny pushing back | Why |
|---|---|---|---|---|---|---|
| Prompt Engineering Longevity and Artificial Social Intelligence | 4 | 5 | 2 | 1 | Lenny frames the debate around prompt engineering relevance by quoting Reid Hoffman. Sander introduces the concept of Artificial Social Intelligence and cites a 70% accuracy boost in medical coding through detailed prompt engineering. | |
| Conversational Prompting Versus Product-Focused Prompt Engineering | 5 | 3 | 1 | 1 | Lenny introduces concrete product examples like Granola, Bolt, Lovable, and v0 to articulate the difference between conversational and embedded production prompts. Sander adopts Lenny's phrasing of product-focused prompt engineering. | |
| Few-Shot Prompting and Optimal Prompt Formatting | 5 | 6 | 1 | 1 | Lenny confuses zero-shot and one-shot prompting before Sander gently clarifies standard ML indexing. Lenny then demonstrates his knowledge by citing Y Combinator discussions on RLHF XML post-training formats. | |
| Debunking Role Prompting and Emotional Manipulation Myths | 4 | 7 | 3 | 2 | Sander dismantles the popular belief that role prompting improves accuracy on reasoning tasks, citing benchmark data with negligible 0.01 variance. Lenny acknowledges his own heavy reliance on copywriter role prompts and switches his assumptions. | |
| Decomposition Strategies and Self-Criticism Prompting | 4 | 6 | 1 | 1 | Lenny asks how decomposition differs from chain-of-thought prompting. Sander illustrates with an agentic car dealership return policy example, walking through sub-problem isolation before running tool calls. | |
| Context Architecture, Suicidal Ideation Research, and Prompt Caching | 5 | 6 | 2 | 1 | Lenny neatly recaps the first four prompting strategies and shares his own workflow using Claude for guest preparation. Sander explains NLP entrapment research and the cost/latency benefits of top-of-prompt caching. | |
| Advanced Prompting: Ensembling and Mixture of Reasoning Experts | 4 | 6 | 1 | 1 | Sander explains multi-agent ensembling and the Mixture of Reasoning Experts technique developed at Stanford. Lenny asks a clarifying question about single-model vs multi-model ensembling setups. | |
| Sponsor Message: Vanta Compliance Automation | 5 | 5 | 1 | 1 | Following a sponsor read with Christina Cacioppo, Lenny probes whether explicit chain-of-thought prompting is obsolete with modern reasoning models. Sander notes edge-case non-reasoning fallbacks in production before sharing a practical reality check on his own conversational shorthand prompts. | |
| AI Red Teaming, Grandmother Jailbreaks, and HackAPrompt | 3 | 7 | 2 | 1 | Sander introduces prompt injection and red teaming, highlighting the grandmother bomb bedtime story jailbreak and his award-winning HackAPrompt competition. He differentiates classical security flaws from looming agentic physical safety concerns. | |
| CBRN Threats, Obfuscation Vectors, and Agentic Code Vulnerabilities | 4 | 8 | 2 | 1 | Sander details dangerous CBRN uplift vectors and illustrates biblical obfuscation via Ender's Game before demonstrating Base64 and Spanish translation jailbreaks. Lenny links the discussion to indirect prompt injection in autonomous coding agents. | |
| Why System Prompts and Guardrails Fail: Mitigating Prompt Injections | 4 | 8 | 4 | 2 | Sander forcefully dismisses system prompt instructions and external guardrails as fundamentally ineffective due to the intelligence gap between guardrail models and base LLMs. He states plainly that prompt injection is an unsolvable problem. | |
| Misalignment Risks, Autonomous SDR Scenarios, and AI Regulation | 5 | 6 | 2 | 2 | Lenny references Asimov's robot laws, Anthropic blackmail test cases, and the paperclip maximizer. Sander explains his conversion to believing misalignment risks through Palisade chess exploits and paints an extreme autonomous SDR scenario. | |
| Lightning Round: Books, Culture, Everyday Tech, and Life Mottos | 5 | 3 | 0 | 0 | In the lightning round, Sander shares his love for The River of Doubt, Black Mirror, and the Daylight Computer DC-1. Lenny surprises Sander by revealing he was an early angel investor in the Daylight Computer. |