Sep 20, 2024 · 1h 7m · latent-space
The Ultimate Guide to Prompting - with Sander Schulhoff from LearnPrompting.org
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
Sander Schulhoff, creator of LearnPrompting.org and lead author of The Prompt Report, joins the Latent Space podcast to unpack the empirical science of prompt engineering, adversarial AI safety, and the programmatic evolution toward AI engineering.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 27.2% of the talking time here. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Sander rejects mainstream assumptions about role prompting, pointing out his empirical benchmark where prompting a model as an idiot outperformed prompting it as a Harvard math professor on MMLU.
Hardest push from the hosts ▶ 43:54 Swyx proves the necessity of cost-aware model cascadesSwyx directly counters Sander's dismissal of model routing by citing specific pricing figures ($5/M tokens vs $0.15/M tokens) to demonstrate why ensembling cheap models under a smart judge is essential for production.
Biggest teaching moment ▶ 50:53 The formal boundary between prompt injection and jailbreakingSander unpacks the widespread confusion in industry and literature between prompt injection (developer instructions overridden by user data) and jailbreaking (user directly bypassing safety guardrails).
The host holds their own ▶ 19:31 Swyx defending multi-persona reasoning and synthetic generationSwyx counters Sander's critique of role prompting by citing concrete research examples including Salesforce's DEI benchmark and Tencent's billion-persona synthetic data paper.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Sander Schulhoff's Background and Journey to The Prompt Report | 3 | 1 | 0 | 0 | The episode opens with standard biographical interviewing where Swyx and Alessio prompt Sander to narrate his progression from Diplomacy and MineRL reinforcement learning to founding LearnPrompting and leading the Hack-A-Prompt and Prompt Report research initiatives. Sander delivers an extensive monologue detailing his academic path and research publications with minimal interruption. | |
| Systematically Reviewing Academic Literature via the PRISMA Process | 6 | 4 | 1 | 1 | Swyx demonstrates strong literacy regarding current AI literature generators like Sakana AI and Omniscience, prompting a discussion on systematic literature reviews. Sander explains the rigorous PRISMA methodology and how his team benchmarked LLMs against human evaluation to filter thousands of prompting papers. | |
| Categorizing Prompt Strategies and Deconstructing Role-Based Prompting | 7 | 5 | 4 | 6 | Sander takes a contrarian stance by declaring that role prompting and emotion prompting do not work for accuracy-based tasks on modern models, citing empirical benchmarks where an 'idiot' prompt beat a 'genius' prompt. Swyx pushes back vigorously, citing the Six Thinking Hats framework, Salesforce's DEI paper, and Tencent's billion-persona synthetic data generation to show where personas still provide reasoning value. | |
| Few-Shot Prompt Design Principles and Pitfalls | 7 | 5 | 1 | 3 | Sander reviews few-shot design parameters like exemplar ordering and in-distribution formatting. Swyx brings authentic production challenges regarding exemplar leakage in production newsletters, prompting Sander to explain literature showing that models prioritize output structure over exemplar label veracity. | |
| Thought Generation, Chain of Thought, and Decomposition Techniques | 7 | 4 | 3 | 6 | Sander shares his AutodiCoT technique and asserts that reasoning prompts like Chain of Thought should theoretically be obsolete due to post-training alignment. Swyx challenges this premise, noting that general-purpose foundation models inherently require explicit conditioning to focus attention on reasoning sub-spaces. | |
| Ensembling, Self-Consistency, and Cost-Aware Model Cascades | 8 | 3 | 3 | 7 | When discussing ensembling and cost, Sander suggests it is generally more practical to pay for top-tier frontier models rather than engineer complex routing cascades. Swyx pushes back strongly by presenting concrete token economics and his multi-tier production workflow where parallel calls to GPT-4o-mini are judged by a frontier model for significant net savings. | |
| Automated Prompt Engineering, DSPy, and the Role of AI Engineers | 7 | 2 | 0 | 1 | Sander admits how DSPy outperformed 20 hours of his manual prompt engineering in 10 minutes. Swyx connects this directly to his 'Rise of the AI Engineer' thesis, emphasizing that production prompting requires software engineering and programmatic optimization rather than standalone prompt craft. | |
| Jailbreaking, Prompt Injections, and Discoveries from Hack-A-Prompt | 7 | 6 | 2 | 2 | Alessio showcases his security background discussing CTFs and DEFCON red-teaming challenges. Sander provides a crisp breakdown differentiating prompt injection (conflicting developer vs user instructions) from jailbreaking (bypassing model safety filters directly), and highlights the context overflow attack discovered in Hack-A-Prompt. | |
| Multimodal Prompting Nuances, Structured Outputs, and Evaluation Pitfalls | 6 | 6 | 1 | 1 | After briefly exchanging notes on music and video generation difficulties, Sander provides a comprehensive critique of flawed LLM evaluation practices, specifically detailing why unanchored Likert scales and raw token log-probabilities yield misleading results in sensitive domains like healthcare. |