May 30, 2025 · 31m · y-combinator

State-Of-The-Art Prompting For AI Agents · Y Combinator

Garry Tan · 13m spoken Diana Hu · 7m spoken Harj Taggar · 5m spoken Jared Friedman · 3m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of The Light Cone Podcast, Y Combinator partners analyze cutting-edge prompt engineering strategies, meta-prompting workflows, and evaluation frameworks used by top AI startups. Drawing on real-world examples and operational playbooks, they demonstrate how treating prompting like human management and embedding deeply with clients enables founders to build scalable, high-defensibility AI agents.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The partners hold 55.3% of the talking time here. How this is scored →

The partners as informed peer 5.7 Guest teaching 3.9 Guest disagreement 0.3 The partners pushing back 0.3
05100:0010:0020:0030:000:57–3:55 · The partners as informed peer 8/10 Analyzing Parahelp's Open-Sourced AI Support Prompt Jared introduces Parahelp's open-sourced prompt and Diana delivers an authoritative, highly technical breakdown of its architecture, role-setting, and XML tag formatting optimized for post-trained RLHF models.3:55–6:01 · The partners as informed peer 7/10 Multi-Tiered Prompt Architecture: System, Developer, and User Roles Gary questions the absence of scenario examples in the prompt, prompting Jared and Diana to clarify that those belong in downstream developer and customer-specific pipeline tiers rather than the core system prompt.6:01–9:05 · The partners as informed peer 6/10 Automated Worked Examples and Introduction to Meta-Prompting Harj highlights automated worked examples, Gary shares insights from YC startup Tropier regarding prompt folding, and Diana details how Jasberry utilizes meta-prompting with hard examples for N+1 bug detection.9:05–12:44 · The partners as informed peer 7/10 Designing LLM Escape Hatches and Production Debug Info Gary explains escape hatches to prevent hallucination, while Jared and Harj counter with their internal YC solution involving a dedicated debug info response parameter acting as a developer to-do list.12:44–17:01 · The partners as informed peer 6/10 Analyzing Thinking Traces and REPL Workflows Hosts discuss Gemini 2.5 thinking traces and REPL workflows, before Gary elaborates deeply on evals being the true company moat derived from on-the-ground ethnography in niche industries.17:01–23:18 · The partners as informed peer 3/10 The Palantir Playbook: Forward-Deployed Engineers as Founders Harj prompts Gary to explain the origin of forward-deployed engineers at Palantir, leading to an extensive historical breakdown by Gary of placing engineers directly with government clients instead of traditional sales reps.23:18–26:12 · The partners as informed peer 7/10 Scaling Vertical AI Agents: GigaML and Happy Robot Diana and Harj share concrete examples of YC startups like GigaML and Happy Robot applying the forward-deployed engineer playbook to close major enterprise contracts through rapid on-site prompt iteration.26:12–29:47 · The partners as informed peer 7/10 Evaluating Model Personalities: O3 vs. Gemini 2.5 Pro Gary introduces investor scoring rubrics, leading Harj and Diana to compare model behaviors, contrasting O3's rigid rule adherence with Gemini 2.5 Pro's flexible, context-aware reasoning.29:47–31:07 · The partners as informed peer 0/10 Conclusion: Prompting as Management and the Kaizen Mindset In a closing monologue, Gary synthesizes prompt engineering with management principles and the Japanese manufacturing concept of Kaizen, without host intervention.0:57–3:55 · Guest teaching 1/10 Analyzing Parahelp's Open-Sourced AI Support Prompt Jared introduces Parahelp's open-sourced prompt and Diana delivers an authoritative, highly technical breakdown of its architecture, role-setting, and XML tag formatting optimized for post-trained RLHF models.3:55–6:01 · Guest teaching 2/10 Multi-Tiered Prompt Architecture: System, Developer, and User Roles Gary questions the absence of scenario examples in the prompt, prompting Jared and Diana to clarify that those belong in downstream developer and customer-specific pipeline tiers rather than the core system prompt.6:01–9:05 · Guest teaching 6/10 Automated Worked Examples and Introduction to Meta-Prompting Harj highlights automated worked examples, Gary shares insights from YC startup Tropier regarding prompt folding, and Diana details how Jasberry utilizes meta-prompting with hard examples for N+1 bug detection.9:05–12:44 · Guest teaching 4/10 Designing LLM Escape Hatches and Production Debug Info Gary explains escape hatches to prevent hallucination, while Jared and Harj counter with their internal YC solution involving a dedicated debug info response parameter acting as a developer to-do list.12:44–17:01 · Guest teaching 7/10 Analyzing Thinking Traces and REPL Workflows Hosts discuss Gemini 2.5 thinking traces and REPL workflows, before Gary elaborates deeply on evals being the true company moat derived from on-the-ground ethnography in niche industries.17:01–23:18 · Guest teaching 8/10 The Palantir Playbook: Forward-Deployed Engineers as Founders Harj prompts Gary to explain the origin of forward-deployed engineers at Palantir, leading to an extensive historical breakdown by Gary of placing engineers directly with government clients instead of traditional sales reps.23:18–26:12 · Guest teaching 1/10 Scaling Vertical AI Agents: GigaML and Happy Robot Diana and Harj share concrete examples of YC startups like GigaML and Happy Robot applying the forward-deployed engineer playbook to close major enterprise contracts through rapid on-site prompt iteration.26:12–29:47 · Guest teaching 2/10 Evaluating Model Personalities: O3 vs. Gemini 2.5 Pro Gary introduces investor scoring rubrics, leading Harj and Diana to compare model behaviors, contrasting O3's rigid rule adherence with Gemini 2.5 Pro's flexible, context-aware reasoning.29:47–31:07 · Guest teaching 4/10 Conclusion: Prompting as Management and the Kaizen Mindset In a closing monologue, Gary synthesizes prompt engineering with management principles and the Japanese manufacturing concept of Kaizen, without host intervention.0:57–3:55 · Guest disagreement 0/10 Analyzing Parahelp's Open-Sourced AI Support Prompt Jared introduces Parahelp's open-sourced prompt and Diana delivers an authoritative, highly technical breakdown of its architecture, role-setting, and XML tag formatting optimized for post-trained RLHF models.3:55–6:01 · Guest disagreement 1/10 Multi-Tiered Prompt Architecture: System, Developer, and User Roles Gary questions the absence of scenario examples in the prompt, prompting Jared and Diana to clarify that those belong in downstream developer and customer-specific pipeline tiers rather than the core system prompt.6:01–9:05 · Guest disagreement 0/10 Automated Worked Examples and Introduction to Meta-Prompting Harj highlights automated worked examples, Gary shares insights from YC startup Tropier regarding prompt folding, and Diana details how Jasberry utilizes meta-prompting with hard examples for N+1 bug detection.9:05–12:44 · Guest disagreement 0/10 Designing LLM Escape Hatches and Production Debug Info Gary explains escape hatches to prevent hallucination, while Jared and Harj counter with their internal YC solution involving a dedicated debug info response parameter acting as a developer to-do list.12:44–17:01 · Guest disagreement 1/10 Analyzing Thinking Traces and REPL Workflows Hosts discuss Gemini 2.5 thinking traces and REPL workflows, before Gary elaborates deeply on evals being the true company moat derived from on-the-ground ethnography in niche industries.17:01–23:18 · Guest disagreement 1/10 The Palantir Playbook: Forward-Deployed Engineers as Founders Harj prompts Gary to explain the origin of forward-deployed engineers at Palantir, leading to an extensive historical breakdown by Gary of placing engineers directly with government clients instead of traditional sales reps.23:18–26:12 · Guest disagreement 0/10 Scaling Vertical AI Agents: GigaML and Happy Robot Diana and Harj share concrete examples of YC startups like GigaML and Happy Robot applying the forward-deployed engineer playbook to close major enterprise contracts through rapid on-site prompt iteration.26:12–29:47 · Guest disagreement 0/10 Evaluating Model Personalities: O3 vs. Gemini 2.5 Pro Gary introduces investor scoring rubrics, leading Harj and Diana to compare model behaviors, contrasting O3's rigid rule adherence with Gemini 2.5 Pro's flexible, context-aware reasoning.29:47–31:07 · Guest disagreement 0/10 Conclusion: Prompting as Management and the Kaizen Mindset In a closing monologue, Gary synthesizes prompt engineering with management principles and the Japanese manufacturing concept of Kaizen, without host intervention.0:57–3:55 · The partners pushing back 0/10 Analyzing Parahelp's Open-Sourced AI Support Prompt Jared introduces Parahelp's open-sourced prompt and Diana delivers an authoritative, highly technical breakdown of its architecture, role-setting, and XML tag formatting optimized for post-trained RLHF models.3:55–6:01 · The partners pushing back 2/10 Multi-Tiered Prompt Architecture: System, Developer, and User Roles Gary questions the absence of scenario examples in the prompt, prompting Jared and Diana to clarify that those belong in downstream developer and customer-specific pipeline tiers rather than the core system prompt.6:01–9:05 · The partners pushing back 0/10 Automated Worked Examples and Introduction to Meta-Prompting Harj highlights automated worked examples, Gary shares insights from YC startup Tropier regarding prompt folding, and Diana details how Jasberry utilizes meta-prompting with hard examples for N+1 bug detection.9:05–12:44 · The partners pushing back 1/10 Designing LLM Escape Hatches and Production Debug Info Gary explains escape hatches to prevent hallucination, while Jared and Harj counter with their internal YC solution involving a dedicated debug info response parameter acting as a developer to-do list.12:44–17:01 · The partners pushing back 0/10 Analyzing Thinking Traces and REPL Workflows Hosts discuss Gemini 2.5 thinking traces and REPL workflows, before Gary elaborates deeply on evals being the true company moat derived from on-the-ground ethnography in niche industries.17:01–23:18 · The partners pushing back 0/10 The Palantir Playbook: Forward-Deployed Engineers as Founders Harj prompts Gary to explain the origin of forward-deployed engineers at Palantir, leading to an extensive historical breakdown by Gary of placing engineers directly with government clients instead of traditional sales reps.23:18–26:12 · The partners pushing back 0/10 Scaling Vertical AI Agents: GigaML and Happy Robot Diana and Harj share concrete examples of YC startups like GigaML and Happy Robot applying the forward-deployed engineer playbook to close major enterprise contracts through rapid on-site prompt iteration.26:12–29:47 · The partners pushing back 0/10 Evaluating Model Personalities: O3 vs. Gemini 2.5 Pro Gary introduces investor scoring rubrics, leading Harj and Diana to compare model behaviors, contrasting O3's rigid rule adherence with Gemini 2.5 Pro's flexible, context-aware reasoning.29:47–31:07 · The partners pushing back 0/10 Conclusion: Prompting as Management and the Kaizen Mindset In a closing monologue, Gary synthesizes prompt engineering with management principles and the Japanese manufacturing concept of Kaizen, without host intervention.

speaking balance: gold is the partners, purple is the guest (3 minute bins)

0:00 · the partners 63.8% · guest 36.2%0:00 · the partners 63.8% · guest 36.2%3:00 · the partners 91.5% · guest 8.5%3:00 · the partners 91.5% · guest 8.5%6:00 · the partners 66.9% · guest 33.1%6:00 · the partners 66.9% · guest 33.1%9:00 · the partners 78.7% · guest 21.3%9:00 · the partners 78.7% · guest 21.3%12:00 · the partners 75.8% · guest 24.2%12:00 · the partners 75.8% · guest 24.2%15:00 · the partners 19.1% · guest 80.9%15:00 · the partners 19.1% · guest 80.9%18:00 · the partners 13.4% · guest 86.6%18:00 · the partners 13.4% · guest 86.6%21:00 · the partners 26.3% · guest 73.7%21:00 · the partners 26.3% · guest 73.7%24:00 · the partners 88.1% · guest 11.9%24:00 · the partners 88.1% · guest 11.9%27:00 · the partners 51.1% · guest 48.9%27:00 · the partners 51.1% · guest 48.9%30:00 · the partners 0% · guest 100%30:00 · the partners 0% · guest 100%
Sharpest disagreement ▶ 3:55 Gary challenges missing prompt sections

Gary points out an apparent omission in the Parahelp prompt, pushing back on its completeness before Jared clarifies the multi-tiered architecture.

Hardest push from the partners ▶ 4:09 Jared clarifies prompt pipelining

Jared directly reframes Gary's critique by explaining that customer-specific examples are intentionally deferred to later stages to prevent custom consulting creep.

Biggest teaching moment ▶ 17:41 Gary breaks down Palantir's FDE philosophy

Gary provides extensive first-hand institutional knowledge on how Palantir replaced traditional sales representatives with engineers sitting directly alongside government agents.

The partners hold their own ▶ 1:48 Diana explains XML tag prompt structure and RLHF alignment

Diana demonstrates deep technical mastery by explaining why modern LLMs respond better to XML tag schemas based on post-training RLHF data formats.

the scores for every segment, with the reasoning behind each
ChapterTopicThe partners as informed peerGuest teachingGuest disagreementThe partners pushing backWhy
Analyzing Parahelp's Open-Sourced AI Support Prompt 8100 Jared introduces Parahelp's open-sourced prompt and Diana delivers an authoritative, highly technical breakdown of its architecture, role-setting, and XML tag formatting optimized for post-trained RLHF models.
Multi-Tiered Prompt Architecture: System, Developer, and User Roles 7212 Gary questions the absence of scenario examples in the prompt, prompting Jared and Diana to clarify that those belong in downstream developer and customer-specific pipeline tiers rather than the core system prompt.
Automated Worked Examples and Introduction to Meta-Prompting 6600 Harj highlights automated worked examples, Gary shares insights from YC startup Tropier regarding prompt folding, and Diana details how Jasberry utilizes meta-prompting with hard examples for N+1 bug detection.
Designing LLM Escape Hatches and Production Debug Info 7401 Gary explains escape hatches to prevent hallucination, while Jared and Harj counter with their internal YC solution involving a dedicated debug info response parameter acting as a developer to-do list.
Analyzing Thinking Traces and REPL Workflows 6710 Hosts discuss Gemini 2.5 thinking traces and REPL workflows, before Gary elaborates deeply on evals being the true company moat derived from on-the-ground ethnography in niche industries.
The Palantir Playbook: Forward-Deployed Engineers as Founders 3810 Harj prompts Gary to explain the origin of forward-deployed engineers at Palantir, leading to an extensive historical breakdown by Gary of placing engineers directly with government clients instead of traditional sales reps.
Scaling Vertical AI Agents: GigaML and Happy Robot 7100 Diana and Harj share concrete examples of YC startups like GigaML and Happy Robot applying the forward-deployed engineer playbook to close major enterprise contracts through rapid on-site prompt iteration.
Evaluating Model Personalities: O3 vs. Gemini 2.5 Pro 7200 Gary introduces investor scoring rubrics, leading Harj and Diana to compare model behaviors, contrasting O3's rigid rule adherence with Gemini 2.5 Pro's flexible, context-aware reasoning.
Conclusion: Prompting as Management and the Kaizen Mindset 0400 In a closing monologue, Gary synthesizes prompt engineering with management principles and the Japanese manufacturing concept of Kaizen, without host intervention.

Statements from this episode (28)

Disclosure
Tan: Y Combinator Surveyed Over a Dozen Frontier AI Companies on Prompting
“We surveyed more than a dozen companies and got their take right from the frontier of building this stuff, the practical tips.”
Garry Tan May 30, 2025 ▶ 0:44
Assertion Supported
Friedman: Parahelp powers customer support for Perplexity, Replit, and Bolt
“They're actually powering the customer support for Perplexity and Replit and Bolt and a bunch of other like top AI companies now.”
Jared Friedman May 30, 2025 ▶ 1:08
Insight
Friedman: Production prompts serve as the core IP for vertical AI companies
“It's like relatively hard to get these prompts for vertical AI agents cause they're kind of like the crown jewels of the IP of these companies.”
Jared Friedman May 30, 2025 ▶ 1:25
Insight
Diana Hu: Best Prompts Outline Reasoning and Provide Examples
“One big thing about the best prompts is they outline how to reason about the task, and then a big thing is giving it an example, and this is what it does.”
Diana Hu May 30, 2025 ▶ 3:21
Insight
Diana Hu: XML Prompts Produce Better Output Due to Post-Training RLHF
“We found that it makes it a lot easier for LLMs to follow, because a lot of elements were post-trained in RLHF with kind of XML type of input, and it turns out to produce better results.”
Diana Hu May 30, 2025 ▶ 3:43
Insight
Friedman: Vertical AI agents risk turning into bespoke consulting businesses
“And so their challenge, like a lot of these agent companies is like, How do you build a general purpose product when every customer like wants, you know, has like slightly different workflows and like preferences has a really interesting thing that I see the v…”
Jared Friedman May 30, 2025 ▶ 4:22
Insight
Diana Hu outlines emerging multi-tier prompt architecture for AI agents
“So there's this concept of defining the prompt in the system prompt, then there's the developer prompt, and then there's the user prompt. So what this mean is the system prompt is basically almost like defining sort of the high level API of how your company op…”
Diana Hu May 30, 2025 ▶ 5:00
Assertion Not checkable as stated
Friedman: Meta-Prompting Is a Consistent Theme Among YC AI Startups
“Meta prompting, which is one of the things we want to talk about, because that's a consistent theme that keeps coming up when we talk to our AI startups.”
Jared Friedman May 30, 2025 ▶ 6:50
Assertion Not checkable as stated
Hu: Even top LLMs struggle to detect N+1 query bugs alone
“It's actually hard for today for even, like, the best LLMs to find those”
Diana Hu May 30, 2025 ▶ 8:24
Insight
Hu: Example-based prompting is the LLM version of test-driven development
“I think this pattern of sometimes when it's too hard to even kind of write a prose around it, let's just give you an example that turns out to work really well, because it helps LLMs to Reason around complicated tasks and steer it better because you can't quit…”
Diana Hu May 30, 2025 ▶ 8:24
Insight
Tan: LLMs need explicit escape hatches to prevent formatted hallucinations
“The model really wants to actually help you so much that if you just tell it, give me back output in this particular format, even if it doesn't quite have the information it needs, it'll actually just tell you what it thinks you want to hear. And it's literall…”
Garry Tan May 30, 2025 ▶ 9:05
Disclosure
Friedman: YC uses a debug info field in agent response schemas
“In the response format. To give it the ability to have part of the response be essentially a complaint to you, the developer that like you have given it confusing or underspecified information and it doesn't know what to do. And then the nice thing about that …”
Jared Friedman May 30, 2025 ▶ 10:00
Insight
Taggar: Prompting LLMs as Expert Engineers Creates an Effective Optimization Loop
“A very simple way to get started with meta-prompting is to follow the same structure of the prompt as to give it a role and make the role be like, you know, you're an expert prompt engineer who gives really, like, detailed great critiques and advice on how to …”
Harj Taggar May 30, 2025 ▶ 10:43
Insight
Hu: Voice AI Startups Meta-Prompt with Frontier Models Before Distillation
“I think that's a common pattern sometimes for companies when they need to get responses from elements, elements in their product a lot quicker. They do the meta-prompting with a bigger, beefier model, any of the, I don't know, hundreds of billions of parameter…”
Diana Hu May 30, 2025 ▶ 11:12
Insight
Taggar: Gemini Pro Excels at Integrating Feedback Notes into Prompts
“One thing I found useful is as you're using it, if you just note down in a Google doc things that you're seeing, just the outputs not being how you want or ways that you can think to improve it, you can just write those in note form and then give Gemini Pro li…”
Harj Taggar May 30, 2025 ▶ 12:18
Assertion Supported
Friedman: Gemini API added thinking traces for debugging prompts
“Cause if you're just using Gemini via the API until recently, you did not get the thinking traces and like the thinking traces are like the critical debugging information to like, understand like What's wrong with your prompt? They just added it to the API. So…”
Jared Friedman May 30, 2025 ▶ 13:00
Insight
Taggar: Gemini Pro's context window enables REPL-style prompt steering
“I think it's an underrated consequence of Gemini pro having such long context windows is you can effectively use it like a rep or go sort of like one by one. I put your prompt on like one example, then literally watch the reasoning trace in real time to figure…”
Harj Taggar May 30, 2025 ▶ 13:19
Insight
Jared Friedman: Evals, not prompts, are the core data asset for AI startups
“Even though we've been saying this for a year or more now, Gary, I think it's still the case that like evals are the true crown jewel Like data asset for all of these companies. Like one reason that power help was willing to open source the prompt is they told…”
Jared Friedman May 30, 2025 ▶ 14:27
Insight
Garry Tan: Vertical AI moats require sitting with domain workers to build evals
“You can't get the evals unless you are sitting literally side by side with people who are doing X, Y, or Z knowledge work. You know, you need to sit next to the tractor sales regional manager and understand, well, you know, this person cares about, you know, t…”
Garry Tan May 30, 2025 ▶ 15:02
Opinion
Tan: Former Palantir FDEs are becoming top YC founders
“Forward deployed engineers who Came up through that system at Palantir now. They're turning out to be some of the best founders at YC, actually.”
Garry Tan May 30, 2025 ▶ 20:25
Insight
Tan: Technical founders must not outsource forward-deployed engineering
“Like, you definitely can't farm this out. Like, literally, the founders themselves, they're technical. They have to be the great product people. They have to be the ethnographer. They have to be the designer. You want the person on the second meeting to see th…”
Garry Tan May 30, 2025 ▶ 22:57
Assertion Not publicly verifiable
Hu: Happy Robot Closed 7-Figure Deals with Top 3 Logistics Brokers
“Happy Robot, who has sold seven figure contracts to the top three largest logistic brokers in the world. They're built AI voice agents for that.”
Diana Hu May 30, 2025 ▶ 25:34
Opinion
Hu: Claude is naturally human-steerable while Llama requires heavy prompting
“One of the things that's known a lot is Claude is sort of the more happy and more human steerable model, and the other one is Lama. Four is one that needs a lot more steering. It's almost like talking to a developer, and part of it could be an artifact of not …”
Diana Hu May 30, 2025 ▶ 26:28
Disclosure
Tan: Y Combinator uses LLMs internally to evaluate prospective venture investors
“Well, one of the things we've been using LLMs for internally is actually helping founders figure out who they should take money from.”
Garry Tan May 30, 2025 ▶ 27:03
Opinion
Taggar: o3 follows rubrics rigidly while Gemini 2.5 Pro reasons flexibly
“Oh, three was very rigid, actually. Like it really sticks to the rubric. It's heavily penalizes for anything that doesn't fit like the rubric that you've given it. Whereas Gemini 2.5 pro was actually quite good at being flexible in that it would apply the rubr…”
Harj Taggar May 30, 2025 ▶ 27:53
Opinion
Tan: Benchmark and Thrive run immaculate processes and never ghost founders
“You know, sometimes you have investors like a benchmark or a Thrive. It's like, yeah, take their money right away. Their process is immaculate. They never ghost anyone. They answer their emails faster than most founders. It's, you know, very impressive.”
Garry Tan May 30, 2025 ▶ 29:02
Insight
Garry Tan: Prompt engineering is like learning how to manage a person
“It also kind of feels like learning how to manage a person. Where it's like, how do I actually communicate you know, the things that they need to know in order to make a good decision? And how do I make sure that they know you know, how I'm going to evaluate a…”
Garry Tan May 30, 2025 ▶ 30:16
Insight
Garry Tan: Meta-prompting is the AI equivalent of Kaizen process improvement
“The people who are the absolute best at improving the process are the people actually doing it, and it's literally why, ah, Japanese cars got so good in the nineties, and that's meta-prompting to me.”
Garry Tan May 30, 2025 ▶ 30:44
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.