Sep 20, 2024 · 1h 7m · latent-space

The Ultimate Guide to Prompting - with Sander Schulhoff from LearnPrompting.org

Sander Schulhoff · 45m spoken Shawn Wang · 11m spoken Alessio Fanelli · 5m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Sander Schulhoff, creator of LearnPrompting.org and lead author of The Prompt Report, joins the Latent Space podcast to unpack the empirical science of prompt engineering, adversarial AI safety, and the programmatic evolution toward AI engineering.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 27.2% of the talking time here. How this is scored →

The hosts as informed peer 6.4 Guest teaching 4.0 Guest disagreement 1.7 The hosts pushing back 3.0
05100:0015:0030:0045:001:00:000:03–8:04 · The hosts as informed peer 3/10 Sander Schulhoff's Background and Journey to The Prompt Report The episode opens with standard biographical interviewing where Swyx and Alessio prompt Sander to narrate his progression from Diplomacy and MineRL reinforcement learning to founding LearnPrompting and leading the Hack-A-Prompt and Prompt Report research initiatives. Sander delivers an extensive monologue detailing his academic path and research publications with minimal interruption.8:05–12:22 · The hosts as informed peer 6/10 Systematically Reviewing Academic Literature via the PRISMA Process Swyx demonstrates strong literacy regarding current AI literature generators like Sakana AI and Omniscience, prompting a discussion on systematic literature reviews. Sander explains the rigorous PRISMA methodology and how his team benchmarked LLMs against human evaluation to filter thousands of prompting papers.12:22–21:35 · The hosts as informed peer 7/10 Categorizing Prompt Strategies and Deconstructing Role-Based Prompting Sander takes a contrarian stance by declaring that role prompting and emotion prompting do not work for accuracy-based tasks on modern models, citing empirical benchmarks where an 'idiot' prompt beat a 'genius' prompt. Swyx pushes back vigorously, citing the Six Thinking Hats framework, Salesforce's DEI paper, and Tencent's billion-persona synthetic data generation to show where personas still provide reasoning value.21:35–28:24 · The hosts as informed peer 7/10 Few-Shot Prompt Design Principles and Pitfalls Sander reviews few-shot design parameters like exemplar ordering and in-distribution formatting. Swyx brings authentic production challenges regarding exemplar leakage in production newsletters, prompting Sander to explain literature showing that models prioritize output structure over exemplar label veracity.28:27–37:38 · The hosts as informed peer 7/10 Thought Generation, Chain of Thought, and Decomposition Techniques Sander shares his AutodiCoT technique and asserts that reasoning prompts like Chain of Thought should theoretically be obsolete due to post-training alignment. Swyx challenges this premise, noting that general-purpose foundation models inherently require explicit conditioning to focus attention on reasoning sub-spaces.37:40–44:48 · The hosts as informed peer 8/10 Ensembling, Self-Consistency, and Cost-Aware Model Cascades When discussing ensembling and cost, Sander suggests it is generally more practical to pay for top-tier frontier models rather than engineer complex routing cascades. Swyx pushes back strongly by presenting concrete token economics and his multi-tier production workflow where parallel calls to GPT-4o-mini are judged by a frontier model for significant net savings.44:49–49:12 · The hosts as informed peer 7/10 Automated Prompt Engineering, DSPy, and the Role of AI Engineers Sander admits how DSPy outperformed 20 hours of his manual prompt engineering in 10 minutes. Swyx connects this directly to his 'Rise of the AI Engineer' thesis, emphasizing that production prompting requires software engineering and programmatic optimization rather than standalone prompt craft.49:13–57:08 · The hosts as informed peer 7/10 Jailbreaking, Prompt Injections, and Discoveries from Hack-A-Prompt Alessio showcases his security background discussing CTFs and DEFCON red-teaming challenges. Sander provides a crisp breakdown differentiating prompt injection (conflicting developer vs user instructions) from jailbreaking (bypassing model safety filters directly), and highlights the context overflow attack discovered in Hack-A-Prompt.57:12–1:04:22 · The hosts as informed peer 6/10 Multimodal Prompting Nuances, Structured Outputs, and Evaluation Pitfalls After briefly exchanging notes on music and video generation difficulties, Sander provides a comprehensive critique of flawed LLM evaluation practices, specifically detailing why unanchored Likert scales and raw token log-probabilities yield misleading results in sensitive domains like healthcare.0:03–8:04 · Guest teaching 1/10 Sander Schulhoff's Background and Journey to The Prompt Report The episode opens with standard biographical interviewing where Swyx and Alessio prompt Sander to narrate his progression from Diplomacy and MineRL reinforcement learning to founding LearnPrompting and leading the Hack-A-Prompt and Prompt Report research initiatives. Sander delivers an extensive monologue detailing his academic path and research publications with minimal interruption.8:05–12:22 · Guest teaching 4/10 Systematically Reviewing Academic Literature via the PRISMA Process Swyx demonstrates strong literacy regarding current AI literature generators like Sakana AI and Omniscience, prompting a discussion on systematic literature reviews. Sander explains the rigorous PRISMA methodology and how his team benchmarked LLMs against human evaluation to filter thousands of prompting papers.12:22–21:35 · Guest teaching 5/10 Categorizing Prompt Strategies and Deconstructing Role-Based Prompting Sander takes a contrarian stance by declaring that role prompting and emotion prompting do not work for accuracy-based tasks on modern models, citing empirical benchmarks where an 'idiot' prompt beat a 'genius' prompt. Swyx pushes back vigorously, citing the Six Thinking Hats framework, Salesforce's DEI paper, and Tencent's billion-persona synthetic data generation to show where personas still provide reasoning value.21:35–28:24 · Guest teaching 5/10 Few-Shot Prompt Design Principles and Pitfalls Sander reviews few-shot design parameters like exemplar ordering and in-distribution formatting. Swyx brings authentic production challenges regarding exemplar leakage in production newsletters, prompting Sander to explain literature showing that models prioritize output structure over exemplar label veracity.28:27–37:38 · Guest teaching 4/10 Thought Generation, Chain of Thought, and Decomposition Techniques Sander shares his AutodiCoT technique and asserts that reasoning prompts like Chain of Thought should theoretically be obsolete due to post-training alignment. Swyx challenges this premise, noting that general-purpose foundation models inherently require explicit conditioning to focus attention on reasoning sub-spaces.37:40–44:48 · Guest teaching 3/10 Ensembling, Self-Consistency, and Cost-Aware Model Cascades When discussing ensembling and cost, Sander suggests it is generally more practical to pay for top-tier frontier models rather than engineer complex routing cascades. Swyx pushes back strongly by presenting concrete token economics and his multi-tier production workflow where parallel calls to GPT-4o-mini are judged by a frontier model for significant net savings.44:49–49:12 · Guest teaching 2/10 Automated Prompt Engineering, DSPy, and the Role of AI Engineers Sander admits how DSPy outperformed 20 hours of his manual prompt engineering in 10 minutes. Swyx connects this directly to his 'Rise of the AI Engineer' thesis, emphasizing that production prompting requires software engineering and programmatic optimization rather than standalone prompt craft.49:13–57:08 · Guest teaching 6/10 Jailbreaking, Prompt Injections, and Discoveries from Hack-A-Prompt Alessio showcases his security background discussing CTFs and DEFCON red-teaming challenges. Sander provides a crisp breakdown differentiating prompt injection (conflicting developer vs user instructions) from jailbreaking (bypassing model safety filters directly), and highlights the context overflow attack discovered in Hack-A-Prompt.57:12–1:04:22 · Guest teaching 6/10 Multimodal Prompting Nuances, Structured Outputs, and Evaluation Pitfalls After briefly exchanging notes on music and video generation difficulties, Sander provides a comprehensive critique of flawed LLM evaluation practices, specifically detailing why unanchored Likert scales and raw token log-probabilities yield misleading results in sensitive domains like healthcare.0:03–8:04 · Guest disagreement 0/10 Sander Schulhoff's Background and Journey to The Prompt Report The episode opens with standard biographical interviewing where Swyx and Alessio prompt Sander to narrate his progression from Diplomacy and MineRL reinforcement learning to founding LearnPrompting and leading the Hack-A-Prompt and Prompt Report research initiatives. Sander delivers an extensive monologue detailing his academic path and research publications with minimal interruption.8:05–12:22 · Guest disagreement 1/10 Systematically Reviewing Academic Literature via the PRISMA Process Swyx demonstrates strong literacy regarding current AI literature generators like Sakana AI and Omniscience, prompting a discussion on systematic literature reviews. Sander explains the rigorous PRISMA methodology and how his team benchmarked LLMs against human evaluation to filter thousands of prompting papers.12:22–21:35 · Guest disagreement 4/10 Categorizing Prompt Strategies and Deconstructing Role-Based Prompting Sander takes a contrarian stance by declaring that role prompting and emotion prompting do not work for accuracy-based tasks on modern models, citing empirical benchmarks where an 'idiot' prompt beat a 'genius' prompt. Swyx pushes back vigorously, citing the Six Thinking Hats framework, Salesforce's DEI paper, and Tencent's billion-persona synthetic data generation to show where personas still provide reasoning value.21:35–28:24 · Guest disagreement 1/10 Few-Shot Prompt Design Principles and Pitfalls Sander reviews few-shot design parameters like exemplar ordering and in-distribution formatting. Swyx brings authentic production challenges regarding exemplar leakage in production newsletters, prompting Sander to explain literature showing that models prioritize output structure over exemplar label veracity.28:27–37:38 · Guest disagreement 3/10 Thought Generation, Chain of Thought, and Decomposition Techniques Sander shares his AutodiCoT technique and asserts that reasoning prompts like Chain of Thought should theoretically be obsolete due to post-training alignment. Swyx challenges this premise, noting that general-purpose foundation models inherently require explicit conditioning to focus attention on reasoning sub-spaces.37:40–44:48 · Guest disagreement 3/10 Ensembling, Self-Consistency, and Cost-Aware Model Cascades When discussing ensembling and cost, Sander suggests it is generally more practical to pay for top-tier frontier models rather than engineer complex routing cascades. Swyx pushes back strongly by presenting concrete token economics and his multi-tier production workflow where parallel calls to GPT-4o-mini are judged by a frontier model for significant net savings.44:49–49:12 · Guest disagreement 0/10 Automated Prompt Engineering, DSPy, and the Role of AI Engineers Sander admits how DSPy outperformed 20 hours of his manual prompt engineering in 10 minutes. Swyx connects this directly to his 'Rise of the AI Engineer' thesis, emphasizing that production prompting requires software engineering and programmatic optimization rather than standalone prompt craft.49:13–57:08 · Guest disagreement 2/10 Jailbreaking, Prompt Injections, and Discoveries from Hack-A-Prompt Alessio showcases his security background discussing CTFs and DEFCON red-teaming challenges. Sander provides a crisp breakdown differentiating prompt injection (conflicting developer vs user instructions) from jailbreaking (bypassing model safety filters directly), and highlights the context overflow attack discovered in Hack-A-Prompt.57:12–1:04:22 · Guest disagreement 1/10 Multimodal Prompting Nuances, Structured Outputs, and Evaluation Pitfalls After briefly exchanging notes on music and video generation difficulties, Sander provides a comprehensive critique of flawed LLM evaluation practices, specifically detailing why unanchored Likert scales and raw token log-probabilities yield misleading results in sensitive domains like healthcare.0:03–8:04 · The hosts pushing back 0/10 Sander Schulhoff's Background and Journey to The Prompt Report The episode opens with standard biographical interviewing where Swyx and Alessio prompt Sander to narrate his progression from Diplomacy and MineRL reinforcement learning to founding LearnPrompting and leading the Hack-A-Prompt and Prompt Report research initiatives. Sander delivers an extensive monologue detailing his academic path and research publications with minimal interruption.8:05–12:22 · The hosts pushing back 1/10 Systematically Reviewing Academic Literature via the PRISMA Process Swyx demonstrates strong literacy regarding current AI literature generators like Sakana AI and Omniscience, prompting a discussion on systematic literature reviews. Sander explains the rigorous PRISMA methodology and how his team benchmarked LLMs against human evaluation to filter thousands of prompting papers.12:22–21:35 · The hosts pushing back 6/10 Categorizing Prompt Strategies and Deconstructing Role-Based Prompting Sander takes a contrarian stance by declaring that role prompting and emotion prompting do not work for accuracy-based tasks on modern models, citing empirical benchmarks where an 'idiot' prompt beat a 'genius' prompt. Swyx pushes back vigorously, citing the Six Thinking Hats framework, Salesforce's DEI paper, and Tencent's billion-persona synthetic data generation to show where personas still provide reasoning value.21:35–28:24 · The hosts pushing back 3/10 Few-Shot Prompt Design Principles and Pitfalls Sander reviews few-shot design parameters like exemplar ordering and in-distribution formatting. Swyx brings authentic production challenges regarding exemplar leakage in production newsletters, prompting Sander to explain literature showing that models prioritize output structure over exemplar label veracity.28:27–37:38 · The hosts pushing back 6/10 Thought Generation, Chain of Thought, and Decomposition Techniques Sander shares his AutodiCoT technique and asserts that reasoning prompts like Chain of Thought should theoretically be obsolete due to post-training alignment. Swyx challenges this premise, noting that general-purpose foundation models inherently require explicit conditioning to focus attention on reasoning sub-spaces.37:40–44:48 · The hosts pushing back 7/10 Ensembling, Self-Consistency, and Cost-Aware Model Cascades When discussing ensembling and cost, Sander suggests it is generally more practical to pay for top-tier frontier models rather than engineer complex routing cascades. Swyx pushes back strongly by presenting concrete token economics and his multi-tier production workflow where parallel calls to GPT-4o-mini are judged by a frontier model for significant net savings.44:49–49:12 · The hosts pushing back 1/10 Automated Prompt Engineering, DSPy, and the Role of AI Engineers Sander admits how DSPy outperformed 20 hours of his manual prompt engineering in 10 minutes. Swyx connects this directly to his 'Rise of the AI Engineer' thesis, emphasizing that production prompting requires software engineering and programmatic optimization rather than standalone prompt craft.49:13–57:08 · The hosts pushing back 2/10 Jailbreaking, Prompt Injections, and Discoveries from Hack-A-Prompt Alessio showcases his security background discussing CTFs and DEFCON red-teaming challenges. Sander provides a crisp breakdown differentiating prompt injection (conflicting developer vs user instructions) from jailbreaking (bypassing model safety filters directly), and highlights the context overflow attack discovered in Hack-A-Prompt.57:12–1:04:22 · The hosts pushing back 1/10 Multimodal Prompting Nuances, Structured Outputs, and Evaluation Pitfalls After briefly exchanging notes on music and video generation difficulties, Sander provides a comprehensive critique of flawed LLM evaluation practices, specifically detailing why unanchored Likert scales and raw token log-probabilities yield misleading results in sensitive domains like healthcare.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 19.1% · guest 80.9%0:00 · the hosts 19.1% · guest 80.9%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 22% · guest 78%6:00 · the hosts 22% · guest 78%9:00 · the hosts 24% · guest 76%9:00 · the hosts 24% · guest 76%12:00 · the hosts 13.2% · guest 86.8%12:00 · the hosts 13.2% · guest 86.8%15:00 · the hosts 19.6% · guest 80.4%15:00 · the hosts 19.6% · guest 80.4%18:00 · the hosts 50.2% · guest 49.8%18:00 · the hosts 50.2% · guest 49.8%21:00 · the hosts 19.4% · guest 80.6%21:00 · the hosts 19.4% · guest 80.6%24:00 · the hosts 42.5% · guest 57.5%24:00 · the hosts 42.5% · guest 57.5%27:00 · the hosts 44% · guest 56%27:00 · the hosts 44% · guest 56%30:00 · the hosts 7.8% · guest 92.2%30:00 · the hosts 7.8% · guest 92.2%33:00 · the hosts 55.8% · guest 44.2%33:00 · the hosts 55.8% · guest 44.2%36:00 · the hosts 21.6% · guest 78.4%36:00 · the hosts 21.6% · guest 78.4%39:00 · the hosts 28.1% · guest 71.9%39:00 · the hosts 28.1% · guest 71.9%42:00 · the hosts 35.4% · guest 64.6%42:00 · the hosts 35.4% · guest 64.6%45:00 · the hosts 48.4% · guest 51.6%45:00 · the hosts 48.4% · guest 51.6%48:00 · the hosts 70.1% · guest 29.9%48:00 · the hosts 70.1% · guest 29.9%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 20.9% · guest 79.1%54:00 · the hosts 20.9% · guest 79.1%57:00 · the hosts 26.1% · guest 73.9%57:00 · the hosts 26.1% · guest 73.9%1:00:00 · the hosts 19.4% · guest 80.6%1:00:00 · the hosts 19.4% · guest 80.6%1:03:00 · the hosts 11.2% · guest 88.8%1:03:00 · the hosts 11.2% · guest 88.8%1:06:00 · the hosts 30.4% · guest 69.6%1:06:00 · the hosts 30.4% · guest 69.6%
Sharpest disagreement ▶ 16:40 Dismissing role prompting efficacy for accuracy tasks

Sander rejects mainstream assumptions about role prompting, pointing out his empirical benchmark where prompting a model as an idiot outperformed prompting it as a Harvard math professor on MMLU.

Hardest push from the hosts ▶ 43:54 Swyx proves the necessity of cost-aware model cascades

Swyx directly counters Sander's dismissal of model routing by citing specific pricing figures ($5/M tokens vs $0.15/M tokens) to demonstrate why ensembling cheap models under a smart judge is essential for production.

Biggest teaching moment ▶ 50:53 The formal boundary between prompt injection and jailbreaking

Sander unpacks the widespread confusion in industry and literature between prompt injection (developer instructions overridden by user data) and jailbreaking (user directly bypassing safety guardrails).

The host holds their own ▶ 19:31 Swyx defending multi-persona reasoning and synthetic generation

Swyx counters Sander's critique of role prompting by citing concrete research examples including Salesforce's DEI benchmark and Tencent's billion-persona synthetic data paper.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Sander Schulhoff's Background and Journey to The Prompt Report 3100 The episode opens with standard biographical interviewing where Swyx and Alessio prompt Sander to narrate his progression from Diplomacy and MineRL reinforcement learning to founding LearnPrompting and leading the Hack-A-Prompt and Prompt Report research initiatives. Sander delivers an extensive monologue detailing his academic path and research publications with minimal interruption.
Systematically Reviewing Academic Literature via the PRISMA Process 6411 Swyx demonstrates strong literacy regarding current AI literature generators like Sakana AI and Omniscience, prompting a discussion on systematic literature reviews. Sander explains the rigorous PRISMA methodology and how his team benchmarked LLMs against human evaluation to filter thousands of prompting papers.
Categorizing Prompt Strategies and Deconstructing Role-Based Prompting 7546 Sander takes a contrarian stance by declaring that role prompting and emotion prompting do not work for accuracy-based tasks on modern models, citing empirical benchmarks where an 'idiot' prompt beat a 'genius' prompt. Swyx pushes back vigorously, citing the Six Thinking Hats framework, Salesforce's DEI paper, and Tencent's billion-persona synthetic data generation to show where personas still provide reasoning value.
Few-Shot Prompt Design Principles and Pitfalls 7513 Sander reviews few-shot design parameters like exemplar ordering and in-distribution formatting. Swyx brings authentic production challenges regarding exemplar leakage in production newsletters, prompting Sander to explain literature showing that models prioritize output structure over exemplar label veracity.
Thought Generation, Chain of Thought, and Decomposition Techniques 7436 Sander shares his AutodiCoT technique and asserts that reasoning prompts like Chain of Thought should theoretically be obsolete due to post-training alignment. Swyx challenges this premise, noting that general-purpose foundation models inherently require explicit conditioning to focus attention on reasoning sub-spaces.
Ensembling, Self-Consistency, and Cost-Aware Model Cascades 8337 When discussing ensembling and cost, Sander suggests it is generally more practical to pay for top-tier frontier models rather than engineer complex routing cascades. Swyx pushes back strongly by presenting concrete token economics and his multi-tier production workflow where parallel calls to GPT-4o-mini are judged by a frontier model for significant net savings.
Automated Prompt Engineering, DSPy, and the Role of AI Engineers 7201 Sander admits how DSPy outperformed 20 hours of his manual prompt engineering in 10 minutes. Swyx connects this directly to his 'Rise of the AI Engineer' thesis, emphasizing that production prompting requires software engineering and programmatic optimization rather than standalone prompt craft.
Jailbreaking, Prompt Injections, and Discoveries from Hack-A-Prompt 7622 Alessio showcases his security background discussing CTFs and DEFCON red-teaming challenges. Sander provides a crisp breakdown differentiating prompt injection (conflicting developer vs user instructions) from jailbreaking (bypassing model safety filters directly), and highlights the context overflow attack discovered in Hack-A-Prompt.
Multimodal Prompting Nuances, Structured Outputs, and Evaluation Pitfalls 6611 After briefly exchanging notes on music and video generation difficulties, Sander provides a comprehensive critique of flawed LLM evaluation practices, specifically detailing why unanchored Likert scales and raw token log-probabilities yield misleading results in sensitive domains like healthcare.

Statements from this episode (18)

Assertion Supported
Schulhoff: Preamble Discovered Prompt Injection Before Riley Goodside
“Preamble is the company that first discovered Prompt Injection, even before Riley, and they, like, responsibly disclosed it, kind of, internally to OpenAI”
Sander Schulhoff Sep 20, 2024 ▶ 4:46
Assertion Supported
Schulhoff: arXiv prohibits and removes undisclosed AI-generated papers
“I found AI-generated papers on Archive, and I flagged them to their staff, and they were like, thank you know, we missed these. Wait, Archive takes them down? Yeah. Oh, I didn't know that. You can't post an AI-generated paper there, especially If you don't say…”
Sander Schulhoff Sep 20, 2024 ▶ 8:29
Opinion
Schulhoff: Role Prompting Does Not Improve Accuracy on Modern LLMs
“For accuracy-based tasks, like MMLU, you're trying to solve a math problem, and maybe you tell the AI that it's a math professor, and you expect it to have improved performance. I really don't think that works. I'm quite certain that doesn't work on more moder…”
Sander Schulhoff Sep 20, 2024 ▶ 17:08
Assertion Supported
Schulhoff: Few-Shot Exemplar Order Can Shift Model Accuracy From 0% to 90%
“How you order your exemplars in the prompt is super important. And we've seen this move accuracy from like zero percent to 90%, like Zero to state of the art on some tasks, which is just ridiculous”
Sander Schulhoff Sep 20, 2024 ▶ 22:22
Assertion Supported
Schulhoff: LLMs Rely More on Prompt Structure Than Exemplar Labels
“There are a number of papers which have found that the label of the exemplar doesn't really matter, and the model reads the exemplars and cares more about structure than label.”
Sander Schulhoff Sep 20, 2024 ▶ 26:41
Insight
Swix: Few-Shot Prompting Is Not Always Superior to Zero-Shot Templates
“Few shot is not necessarily better than zero shot is, which is counterintuitive because you're working harder.”
Shawn Wang Sep 20, 2024 ▶ 28:19
Assertion Not checkable as stated
Schulhoff: GPT-4 Fails to Output Reasoning on 1 in 100 to 1,000 Prompts
“I remember I did a lot of experiments with GPT-IV, and especially when you look at it at scale, so I'll run thousands of prompts against it through the API, and I'll see, you know, every one in a hundred, every one in a thousand outputs no reasoning whatsoever…”
Sander Schulhoff Sep 20, 2024 ▶ 33:03
Assertion Supported
Schulhoff: Self-consistency prompting yields diminishing returns on newer LLMs
“When it came out, it seemed to be quite performant, although more recently, I think as the models have improved, the Performance of this technique has dropped, and you can see that in the evals we run near the end of the paper, where we use it, and it doesn't …”
Sander Schulhoff Sep 20, 2024 ▶ 39:11
Insight
Schulhoff: Researchers should pay for top models instead of engineering routing
“For the most part, designing these systems where you're kind of routing to different levels of intelligence is a really time-consuming and difficult task, and, like, it's probably worth it to just use the smart model And pay for it at this point if you're look…”
Sander Schulhoff Sep 20, 2024 ▶ 42:32
Assertion Supported
Swix: Multi-sampling GPT-4o mini before GPT-4o judging yields net savings
“If I call a GP for a mini 10 times and I do a number of drafts or summaries, and then I have four, oh, judge the summaries that actually is net savings and like a good enough savings then running four, oh, on everything, which given the hundreds and thousands …”
Shawn Wang Sep 20, 2024 ▶ 44:23
Assertion Not checkable as stated
Schulhoff: DSPy Beat 20 Hours of Manual Prompt Engineering in 10 Minutes
“And then I spent 20 hours prompt engineering for a task, and Dyspy beat me in 10 minutes, and that's when I changed my mind.”
Sander Schulhoff Sep 20, 2024 ▶ 45:18
Insight
Schulhoff: Automated Prompt Optimization Fails on Open Generation Without Ground Truth
“One limitation, I guess, is that you really need ground truth labels, so it's harder, if not impossible currently, to optimize open generation tasks, so like Writing, writing newsletters, I suppose. It's harder to automatically optimize those”
Sander Schulhoff Sep 20, 2024 ▶ 45:35
Opinion
Schulhoff: Hiring Dedicated Prompt Engineers Makes No Sense for Most Companies
“I have always viewed prompt engineering as a skill that everybody should and will have, rather than a specialized role to hire for. That being said, there are definitely times where you do need just a prompt engineer. I think for AI companies, it's definitely …”
Sander Schulhoff Sep 20, 2024 ▶ 48:11
Insight
Schulhoff: Prompt injection overrides developer instructions; jailbreaking bypasses model directly
“Basically prompt injection is something that occurs when there is developer input, In the prompt, as well as user input in the prompt. So the developer instructions will say to do one thing, the user input will say to do something else. Jailbreaking is when it…”
Sander Schulhoff Sep 20, 2024 ▶ 51:52
Insight
Schulhoff: Open competitions uncover LLM exploits that paid staff never find
“What's really nice about competitions is that there is stuff that you'll just never find Paying people to do a job. And you'll only find it through random brilliant internet people inspired by thousands of people and the community around them all looking at th…”
Sander Schulhoff Sep 20, 2024 ▶ 53:43
Insight
Schulhoff: Prompting frameworks obscure hidden instructions and hurt reproducibility
“There's a lot of invisible prompts at work on a lot of these frameworks. I hate that. So like, you'll have Oh, this function summarizes input. But if you look behind the scenes, it's using some special summarization instruction. And if you don't have visibilit…”
Sander Schulhoff Sep 20, 2024 ▶ 1:00:55
Insight
Schulhoff: LLMs Have Number Biases and Require Explicit Rubrics for Evaluation
“These methods are super problematic because there is an incredible amount of instability in them, in the sense that models are biased towards outputting certain numbers, and you generally shouldn't say things like, output your result as a number on a scale of …”
Sander Schulhoff Sep 20, 2024 ▶ 1:02:55
Disclosure
Schulhoff: Hack-A-Prompt 2 Aims to Award $500,000 for Harmful AI Dataset
“We're looking to raise and then give away a half million dollars in prizes, and we're going to be creating the most harmful data set ever created, in the sense that this year we're going to be asking people to generate, force the models to generate real-world …”
Sander Schulhoff Sep 20, 2024 ▶ 1:04:37
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.