Sep 25, 2025 · 1h 46m · lennys-podcast

Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar

Hamel Husain · 43m spoken Shreya Shankar · 28m spoken Lenny Rachitsky · 20m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this masterclass, AI evaluation experts Hamel Husain and Shreya Shankar demonstrate how product builders can transition from guesswork and 'vibe checks' to systematic, high-ROI error analysis and automated evaluators.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Lenny holds 21.7% of the talking time here. How this is scored →

Lenny as informed peer 3.9 Guest teaching 4.8 Guest disagreement 1.0 Lenny pushing back 0.4
05100:0020:0040:001:00:001:20:001:40:000:00–2:57 · Lenny as informed peer 3/10 Preview: Why AI Evals Are the Hottest Skill Lenny delivers an introductory monologue summarizing the rising importance of AI evals and framing Hamel and Shreya's course on Maven before cutting to sponsor reads.2:58–10:07 · Lenny as informed peer 4/10 Sponsor: Fin AI Agent by Intercom Lenny proposes that evals are simply unit tests for AI, prompting Shreya to gently reframe the concept into a much wider spectrum of data analytics and qualitative discovery.10:07–23:54 · Lenny as informed peer 3/10 Demonstrating Error Analysis and Open Coding on Real Traces Hamel leads a detailed walkthrough of real-world conversational traces from Nurture Boss, instructing Lenny on open coding and capturing upstream error notes manually.23:55–28:09 · Lenny as informed peer 3/10 The Limits of Automation and the Benevolent Dictator Shreya warns against early LLM automation for note-taking due to lack of domain context, and Hamel explains the necessity of appointing a single benevolent dictator over a committee.28:09–39:19 · Lenny as informed peer 4/10 Synthesizing Open Codes into Axial Categories The guests explain theoretical saturation and demonstrate clustering open codes into axial categories with LLMs, referencing foundational machine learning error analysis roots from Andrew Ng.39:21–48:31 · Lenny as informed peer 3/10 Quantifying Error Distributions and Prioritizing Fixes Hamel demonstrates aggregating categorized errors using pivot tables to quantitatively prioritize application fixes before jumping into automated testing.48:31–53:36 · Lenny as informed peer 4/10 Code-Based Evaluators Versus LLM-as-a-Judge Shreya and Hamel distinguish deterministic code evaluators from LLM judges, strongly criticizing 1-to-5 Likert scales in favor of strict binary pass-fail metrics.53:37–1:01:45 · Lenny as informed peer 4/10 Sponsor: Mercury Financial Services for Startups After an ad break, Hamel details aligning LLM judges with human evaluations and warns why overall percentage agreement is misleading without an error confusion matrix.1:01:45–1:09:57 · Lenny as informed peer 6/10 Evals as Dynamic PRDs and the Impact of Criteria Drift Lenny makes an insightful connection between LLM judge prompts and dynamic PRDs, leading Shreya to discuss her research on criteria drift during model evaluation.1:09:58–1:24:12 · Lenny as informed peer 5/10 Debating Vibe Checks, Coding Agents, and A/B Testing The guests dismantle online debates pitting vibes and A/B testing against evals, explaining that coding agents benefit from unique developer-dogfooding dynamics that do not generalize.1:24:13–1:31:56 · Lenny as informed peer 4/10 Common Misconceptions, Custom Tools, and Practical Advice Hamel and Shreya emphasize looking directly at raw traces and show how simple custom internal interfaces maximize iteration velocity and product ROI.1:31:57–1:37:57 · Lenny as informed peer 4/10 Overview of the Maven AI Evals Course and Resources The guests outline their Maven curriculum, supplemental comprehensive handbook, and custom AI assistant bot trained on their entire corpus.1:37:57–1:44:08 · Lenny as informed peer 3/10 Lightning Round: Books, Culture, Tools, and Mottos The interview concludes with a lighthearted lightning round touching on textbooks, fiction, TV shows, favorite developer tools, and mutual admiration between the guests.0:00–2:57 · Guest teaching 1/10 Preview: Why AI Evals Are the Hottest Skill Lenny delivers an introductory monologue summarizing the rising importance of AI evals and framing Hamel and Shreya's course on Maven before cutting to sponsor reads.2:58–10:07 · Guest teaching 5/10 Sponsor: Fin AI Agent by Intercom Lenny proposes that evals are simply unit tests for AI, prompting Shreya to gently reframe the concept into a much wider spectrum of data analytics and qualitative discovery.10:07–23:54 · Guest teaching 6/10 Demonstrating Error Analysis and Open Coding on Real Traces Hamel leads a detailed walkthrough of real-world conversational traces from Nurture Boss, instructing Lenny on open coding and capturing upstream error notes manually.23:55–28:09 · Guest teaching 6/10 The Limits of Automation and the Benevolent Dictator Shreya warns against early LLM automation for note-taking due to lack of domain context, and Hamel explains the necessity of appointing a single benevolent dictator over a committee.28:09–39:19 · Guest teaching 5/10 Synthesizing Open Codes into Axial Categories The guests explain theoretical saturation and demonstrate clustering open codes into axial categories with LLMs, referencing foundational machine learning error analysis roots from Andrew Ng.39:21–48:31 · Guest teaching 5/10 Quantifying Error Distributions and Prioritizing Fixes Hamel demonstrates aggregating categorized errors using pivot tables to quantitatively prioritize application fixes before jumping into automated testing.48:31–53:36 · Guest teaching 6/10 Code-Based Evaluators Versus LLM-as-a-Judge Shreya and Hamel distinguish deterministic code evaluators from LLM judges, strongly criticizing 1-to-5 Likert scales in favor of strict binary pass-fail metrics.53:37–1:01:45 · Guest teaching 6/10 Sponsor: Mercury Financial Services for Startups After an ad break, Hamel details aligning LLM judges with human evaluations and warns why overall percentage agreement is misleading without an error confusion matrix.1:01:45–1:09:57 · Guest teaching 5/10 Evals as Dynamic PRDs and the Impact of Criteria Drift Lenny makes an insightful connection between LLM judge prompts and dynamic PRDs, leading Shreya to discuss her research on criteria drift during model evaluation.1:09:58–1:24:12 · Guest teaching 6/10 Debating Vibe Checks, Coding Agents, and A/B Testing The guests dismantle online debates pitting vibes and A/B testing against evals, explaining that coding agents benefit from unique developer-dogfooding dynamics that do not generalize.1:24:13–1:31:56 · Guest teaching 5/10 Common Misconceptions, Custom Tools, and Practical Advice Hamel and Shreya emphasize looking directly at raw traces and show how simple custom internal interfaces maximize iteration velocity and product ROI.1:31:57–1:37:57 · Guest teaching 4/10 Overview of the Maven AI Evals Course and Resources The guests outline their Maven curriculum, supplemental comprehensive handbook, and custom AI assistant bot trained on their entire corpus.1:37:57–1:44:08 · Guest teaching 3/10 Lightning Round: Books, Culture, Tools, and Mottos The interview concludes with a lighthearted lightning round touching on textbooks, fiction, TV shows, favorite developer tools, and mutual admiration between the guests.0:00–2:57 · Guest disagreement 0/10 Preview: Why AI Evals Are the Hottest Skill Lenny delivers an introductory monologue summarizing the rising importance of AI evals and framing Hamel and Shreya's course on Maven before cutting to sponsor reads.2:58–10:07 · Guest disagreement 1/10 Sponsor: Fin AI Agent by Intercom Lenny proposes that evals are simply unit tests for AI, prompting Shreya to gently reframe the concept into a much wider spectrum of data analytics and qualitative discovery.10:07–23:54 · Guest disagreement 1/10 Demonstrating Error Analysis and Open Coding on Real Traces Hamel leads a detailed walkthrough of real-world conversational traces from Nurture Boss, instructing Lenny on open coding and capturing upstream error notes manually.23:55–28:09 · Guest disagreement 2/10 The Limits of Automation and the Benevolent Dictator Shreya warns against early LLM automation for note-taking due to lack of domain context, and Hamel explains the necessity of appointing a single benevolent dictator over a committee.28:09–39:19 · Guest disagreement 1/10 Synthesizing Open Codes into Axial Categories The guests explain theoretical saturation and demonstrate clustering open codes into axial categories with LLMs, referencing foundational machine learning error analysis roots from Andrew Ng.39:21–48:31 · Guest disagreement 0/10 Quantifying Error Distributions and Prioritizing Fixes Hamel demonstrates aggregating categorized errors using pivot tables to quantitatively prioritize application fixes before jumping into automated testing.48:31–53:36 · Guest disagreement 2/10 Code-Based Evaluators Versus LLM-as-a-Judge Shreya and Hamel distinguish deterministic code evaluators from LLM judges, strongly criticizing 1-to-5 Likert scales in favor of strict binary pass-fail metrics.53:37–1:01:45 · Guest disagreement 1/10 Sponsor: Mercury Financial Services for Startups After an ad break, Hamel details aligning LLM judges with human evaluations and warns why overall percentage agreement is misleading without an error confusion matrix.1:01:45–1:09:57 · Guest disagreement 1/10 Evals as Dynamic PRDs and the Impact of Criteria Drift Lenny makes an insightful connection between LLM judge prompts and dynamic PRDs, leading Shreya to discuss her research on criteria drift during model evaluation.1:09:58–1:24:12 · Guest disagreement 3/10 Debating Vibe Checks, Coding Agents, and A/B Testing The guests dismantle online debates pitting vibes and A/B testing against evals, explaining that coding agents benefit from unique developer-dogfooding dynamics that do not generalize.1:24:13–1:31:56 · Guest disagreement 1/10 Common Misconceptions, Custom Tools, and Practical Advice Hamel and Shreya emphasize looking directly at raw traces and show how simple custom internal interfaces maximize iteration velocity and product ROI.1:31:57–1:37:57 · Guest disagreement 0/10 Overview of the Maven AI Evals Course and Resources The guests outline their Maven curriculum, supplemental comprehensive handbook, and custom AI assistant bot trained on their entire corpus.1:37:57–1:44:08 · Guest disagreement 0/10 Lightning Round: Books, Culture, Tools, and Mottos The interview concludes with a lighthearted lightning round touching on textbooks, fiction, TV shows, favorite developer tools, and mutual admiration between the guests.0:00–2:57 · Lenny pushing back 0/10 Preview: Why AI Evals Are the Hottest Skill Lenny delivers an introductory monologue summarizing the rising importance of AI evals and framing Hamel and Shreya's course on Maven before cutting to sponsor reads.2:58–10:07 · Lenny pushing back 1/10 Sponsor: Fin AI Agent by Intercom Lenny proposes that evals are simply unit tests for AI, prompting Shreya to gently reframe the concept into a much wider spectrum of data analytics and qualitative discovery.10:07–23:54 · Lenny pushing back 0/10 Demonstrating Error Analysis and Open Coding on Real Traces Hamel leads a detailed walkthrough of real-world conversational traces from Nurture Boss, instructing Lenny on open coding and capturing upstream error notes manually.23:55–28:09 · Lenny pushing back 1/10 The Limits of Automation and the Benevolent Dictator Shreya warns against early LLM automation for note-taking due to lack of domain context, and Hamel explains the necessity of appointing a single benevolent dictator over a committee.28:09–39:19 · Lenny pushing back 0/10 Synthesizing Open Codes into Axial Categories The guests explain theoretical saturation and demonstrate clustering open codes into axial categories with LLMs, referencing foundational machine learning error analysis roots from Andrew Ng.39:21–48:31 · Lenny pushing back 0/10 Quantifying Error Distributions and Prioritizing Fixes Hamel demonstrates aggregating categorized errors using pivot tables to quantitatively prioritize application fixes before jumping into automated testing.48:31–53:36 · Lenny pushing back 0/10 Code-Based Evaluators Versus LLM-as-a-Judge Shreya and Hamel distinguish deterministic code evaluators from LLM judges, strongly criticizing 1-to-5 Likert scales in favor of strict binary pass-fail metrics.53:37–1:01:45 · Lenny pushing back 0/10 Sponsor: Mercury Financial Services for Startups After an ad break, Hamel details aligning LLM judges with human evaluations and warns why overall percentage agreement is misleading without an error confusion matrix.1:01:45–1:09:57 · Lenny pushing back 1/10 Evals as Dynamic PRDs and the Impact of Criteria Drift Lenny makes an insightful connection between LLM judge prompts and dynamic PRDs, leading Shreya to discuss her research on criteria drift during model evaluation.1:09:58–1:24:12 · Lenny pushing back 2/10 Debating Vibe Checks, Coding Agents, and A/B Testing The guests dismantle online debates pitting vibes and A/B testing against evals, explaining that coding agents benefit from unique developer-dogfooding dynamics that do not generalize.1:24:13–1:31:56 · Lenny pushing back 0/10 Common Misconceptions, Custom Tools, and Practical Advice Hamel and Shreya emphasize looking directly at raw traces and show how simple custom internal interfaces maximize iteration velocity and product ROI.1:31:57–1:37:57 · Lenny pushing back 0/10 Overview of the Maven AI Evals Course and Resources The guests outline their Maven curriculum, supplemental comprehensive handbook, and custom AI assistant bot trained on their entire corpus.1:37:57–1:44:08 · Lenny pushing back 0/10 Lightning Round: Books, Culture, Tools, and Mottos The interview concludes with a lighthearted lightning round touching on textbooks, fiction, TV shows, favorite developer tools, and mutual admiration between the guests.

speaking balance: gold is Lenny, purple is the guest (3 minute bins)

0:00 · Lenny 74.9% · guest 25.1%0:00 · Lenny 74.9% · guest 25.1%3:00 · Lenny 92.7% · guest 7.3%3:00 · Lenny 92.7% · guest 7.3%6:00 · Lenny 34.7% · guest 65.3%6:00 · Lenny 34.7% · guest 65.3%9:00 · Lenny 5.5% · guest 94.5%9:00 · Lenny 5.5% · guest 94.5%12:00 · Lenny 1.1% · guest 98.9%12:00 · Lenny 1.1% · guest 98.9%15:00 · Lenny 10.3% · guest 89.7%15:00 · Lenny 10.3% · guest 89.7%18:00 · Lenny 8.6% · guest 91.4%18:00 · Lenny 8.6% · guest 91.4%21:00 · Lenny 5.7% · guest 94.3%21:00 · Lenny 5.7% · guest 94.3%24:00 · Lenny 11.2% · guest 88.8%24:00 · Lenny 11.2% · guest 88.8%27:00 · Lenny 14.3% · guest 85.7%27:00 · Lenny 14.3% · guest 85.7%30:00 · Lenny 8.8% · guest 91.2%30:00 · Lenny 8.8% · guest 91.2%33:00 · Lenny 10.9% · guest 89.1%33:00 · Lenny 10.9% · guest 89.1%36:00 · Lenny 19.3% · guest 80.7%36:00 · Lenny 19.3% · guest 80.7%39:00 · Lenny 18% · guest 82%39:00 · Lenny 18% · guest 82%42:00 · Lenny 25.4% · guest 74.6%42:00 · Lenny 25.4% · guest 74.6%45:00 · Lenny 0% · guest 100%45:00 · Lenny 0% · guest 100%48:00 · Lenny 37% · guest 63%48:00 · Lenny 37% · guest 63%51:00 · Lenny 34% · guest 66%51:00 · Lenny 34% · guest 66%54:00 · Lenny 25% · guest 75%54:00 · Lenny 25% · guest 75%57:00 · Lenny 1% · guest 99%57:00 · Lenny 1% · guest 99%1:00:00 · Lenny 43.1% · guest 56.9%1:00:00 · Lenny 43.1% · guest 56.9%1:03:00 · Lenny 23.8% · guest 76.2%1:03:00 · Lenny 23.8% · guest 76.2%1:06:00 · Lenny 11.1% · guest 88.9%1:06:00 · Lenny 11.1% · guest 88.9%1:09:00 · Lenny 25.1% · guest 74.9%1:09:00 · Lenny 25.1% · guest 74.9%1:12:00 · Lenny 10.9% · guest 89.1%1:12:00 · Lenny 10.9% · guest 89.1%1:15:00 · Lenny 32.8% · guest 67.2%1:15:00 · Lenny 32.8% · guest 67.2%1:18:00 · Lenny 11.7% · guest 88.3%1:18:00 · Lenny 11.7% · guest 88.3%1:21:00 · Lenny 18% · guest 82%1:21:00 · Lenny 18% · guest 82%1:24:00 · Lenny 21.2% · guest 78.8%1:24:00 · Lenny 21.2% · guest 78.8%1:27:00 · Lenny 6.2% · guest 93.8%1:27:00 · Lenny 6.2% · guest 93.8%1:30:00 · Lenny 21% · guest 79%1:30:00 · Lenny 21% · guest 79%1:33:00 · Lenny 14.7% · guest 85.3%1:33:00 · Lenny 14.7% · guest 85.3%1:36:00 · Lenny 28.9% · guest 71.1%1:36:00 · Lenny 28.9% · guest 71.1%1:39:00 · Lenny 17.3% · guest 82.7%1:39:00 · Lenny 17.3% · guest 82.7%1:42:00 · Lenny 31.4% · guest 68.6%1:42:00 · Lenny 31.4% · guest 68.6%1:45:00 · Lenny 35.8% · guest 64.2%1:45:00 · Lenny 35.8% · guest 64.2%
Sharpest disagreement ▶ 1:13:13 Debunking the 'No Evals, Just Vibes' Coding Agent Myth

Shreya and Hamel push back strongly against the viral narrative that top products like Claude Code operate solely on vibe checks without rigorous foundational evals and telemetry.

Hardest push from Lenny ▶ 8:29 Lenny Testing the Unit Test Metaphor

Lenny presses on whether evals can simply be thought of as standard software unit tests, challenging the guests to clarify the exact functional distinction.

Biggest teaching moment ▶ 58:20 Exposing the Trap of High Percentage Agreement in Evals

Hamel demonstrates why a product manager celebrating a 90% judge agreement metric might actually be masking massive failure rates on rare edge cases without a confusion matrix.

Lenny holds their own ▶ 1:00:56 Connecting Automated Judges Directly to Product Requirements Documents

Lenny synthesizes the guests' complex prompt engineering mechanics into the core product management discipline, demonstrating how judge prompts function as continuously executing dynamic PRDs.

the scores for every segment, with the reasoning behind each
ChapterTopicLenny as informed peerGuest teachingGuest disagreementLenny pushing backWhy
Preview: Why AI Evals Are the Hottest Skill 3100 Lenny delivers an introductory monologue summarizing the rising importance of AI evals and framing Hamel and Shreya's course on Maven before cutting to sponsor reads.
Sponsor: Fin AI Agent by Intercom 4511 Lenny proposes that evals are simply unit tests for AI, prompting Shreya to gently reframe the concept into a much wider spectrum of data analytics and qualitative discovery.
Demonstrating Error Analysis and Open Coding on Real Traces 3610 Hamel leads a detailed walkthrough of real-world conversational traces from Nurture Boss, instructing Lenny on open coding and capturing upstream error notes manually.
The Limits of Automation and the Benevolent Dictator 3621 Shreya warns against early LLM automation for note-taking due to lack of domain context, and Hamel explains the necessity of appointing a single benevolent dictator over a committee.
Synthesizing Open Codes into Axial Categories 4510 The guests explain theoretical saturation and demonstrate clustering open codes into axial categories with LLMs, referencing foundational machine learning error analysis roots from Andrew Ng.
Quantifying Error Distributions and Prioritizing Fixes 3500 Hamel demonstrates aggregating categorized errors using pivot tables to quantitatively prioritize application fixes before jumping into automated testing.
Code-Based Evaluators Versus LLM-as-a-Judge 4620 Shreya and Hamel distinguish deterministic code evaluators from LLM judges, strongly criticizing 1-to-5 Likert scales in favor of strict binary pass-fail metrics.
Sponsor: Mercury Financial Services for Startups 4610 After an ad break, Hamel details aligning LLM judges with human evaluations and warns why overall percentage agreement is misleading without an error confusion matrix.
Evals as Dynamic PRDs and the Impact of Criteria Drift 6511 Lenny makes an insightful connection between LLM judge prompts and dynamic PRDs, leading Shreya to discuss her research on criteria drift during model evaluation.
Debating Vibe Checks, Coding Agents, and A/B Testing 5632 The guests dismantle online debates pitting vibes and A/B testing against evals, explaining that coding agents benefit from unique developer-dogfooding dynamics that do not generalize.
Common Misconceptions, Custom Tools, and Practical Advice 4510 Hamel and Shreya emphasize looking directly at raw traces and show how simple custom internal interfaces maximize iteration velocity and product ROI.
Overview of the Maven AI Evals Course and Resources 4400 The guests outline their Maven curriculum, supplemental comprehensive handbook, and custom AI assistant bot trained on their entire corpus.
Lightning Round: Books, Culture, Tools, and Mottos 3300 The interview concludes with a lighthearted lightning round touching on textbooks, fiction, TV shows, favorite developer tools, and mutual admiration between the guests.

Statements from this episode (21)

Opinion
Husain: Product managers, not developers, must lead AI trace error analysis
“Product people have to be in the room and they have to be involved in sort of doing this. You know, usually a developer is not suited to do this, especially if it's not a coding application.”
Hamel Husain Sep 25, 2025 ▶ 17:33
Insight
Husain: Log only the single most upstream error per trace
“Just write down the first thing that you see that's wrong. The most upstream error. Don't worry about all the errors. Just capture the most, the first thing that you see that's wrong and stop and move on.”
Hamel Husain Sep 25, 2025 ▶ 22:14
Insight
Shankar: LLMs fail at initial error analysis due to missing context
“What we usually find when we try to ask an LLM to do this error analysis is it just says the trace looks good because it doesn't have the context needed to understand whether something might be, you know, bad product smell or, you know, not”
Shreya Shankar Sep 25, 2025 ▶ 24:05
Insight
Husain: Appoint a single domain expert for open coding, not committees
“Benevolent dictator is just a catchy term for the fact that when you're doing this open coding, a lot of teens get bogged down in having a committee do this. And for a lot of situations, that's wholly unnecessary. Like, You know, people get really uncomfortabl…”
Hamel Husain Sep 25, 2025 ▶ 25:44
Insight
Husain: LLM-as-a-judge evaluators must use binary scores instead of 1-5 scales
“When you go to building an LLM as a judge, you need a binary score. You don't want to think about, is this like a one, two, three, four, five, like assign a score to it. You can't, that's going to slow it down.”
Hamel Husain Sep 25, 2025 ▶ 26:45
Insight
Husain: Prompting LLMs with 'Axial Codes' Shortcuts Error Categorization
“LLMs know what open codes are, and they know what axial codes are, because it is a concept that's been around for a really long time. So those words help me shortcut, like, what I'm trying to do.”
Hamel Husain Sep 25, 2025 ▶ 33:31
Insight
Shankar: Open code notes must be detailed for LLMs to categorize errors
“This also drives home the point that your open codes have to be detailed, right? You can't just say janky because if the AI is reading janky, it's not going to be able to categorize it. Even a human wouldn't, right? It would have to go and remember why you sai…”
Shreya Shankar Sep 25, 2025 ▶ 42:26
Insight
Husain: Jumping straight to evals without error analysis derails AI products
“You want to usually ground yourself in your actual errors. You don't want to skip this step. And so the reason I'm kind of spending so much time on this is like, this is where people get lost. They go straight into evals. Like, let me just write some tests. An…”
Hamel Husain Sep 25, 2025 ▶ 47:00
Insight
Husain: Prioritize code-based evals over LLM judges to save cost and complexity
“So there's different kinds of evals. One is code-based, which you should try to do if you can, because they're cheaper. You don't have to, you know, LLM as a judge is something, it's like a meta eval. You have to eval that eval to make sure the LLM that's judg…”
Hamel Husain Sep 25, 2025 ▶ 48:05
Insight
Shankar: LLM judges are reliable when scoped to binary failure modes
“People always think like, oh, this is at least as hard as my problem of creating the original agent, and it's not because you're asking the judge to do one thing, evaluate one failure mode. So the scope of the problem is very small and the output of this LLM j…”
Shreya Shankar Sep 25, 2025 ▶ 50:50
Insight
Husain: Raw human-judge agreement is a misleading metric for AI evals
“Now, one thing you should know as a product manager is a lot of people go straight to this, like, agreement. They say, okay, my judge agrees with the human at some percentage of the time. Now that sounds appealing, but it's a very dangerous metric to use becau…”
Hamel Husain Sep 25, 2025 ▶ 58:24
Opinion
Rachitsky: Automated eval judges are the purest form of modern PRDs
“I've had some guests on the podcast recently who've been saying evals are the new PRDs. And if you look at this is exactly what this is like. Product managers, product teams, right? Here's what the product should be. Here's all the requirements. Here's like th…”
Lenny Rachitsky Sep 25, 2025 ▶ 1:01:04
Insight
Shreya Shankar: AI eval rubrics cannot be defined upfront without data
“What's new here is that you can't figure out your rubrics upfront. People's opinions of good and bad change as they review more outputs. They think of failure modes only after seeing 10 outputs they would never have dreamed of in the first place.”
Shreya Shankar Sep 25, 2025 ▶ 1:04:13
Insight
Shreya Shankar: AI products typically need only four to seven LLM evals
“For me, like, between four and seven. It's not that many, because a lot of the failure modes, as Hamill said earlier, can be fixed by just fixing your prompt.”
Shreya Shankar Sep 25, 2025 ▶ 1:05:19
Opinion
Shreya Shankar: AI companies conceal evals because they are competitive moats
“And people don't talk about it because this is their moat, right? So people are not going to go and share all of these things because it makes sense, right? If you are an email writing assistant and you're doing this and you're doing it well, you don't want so…”
Shreya Shankar Sep 25, 2025 ▶ 1:08:23
Insight
Husain: Coding agents differ from other AI products because devs dogfood them
“Coding agents are fundamentally very different than other AI products because the developer is the domain expert. So you can short circuit a lot of things, and also the developer is using it all day long.”
Hamel Husain Sep 25, 2025 ▶ 1:13:53
Insight
Husain: AI evals are just standard data science applied to AI products
“People say the word eval is trying to kind of like carve out this new thing, and saying, you know, evals, and then A-B testing, but if you zoom out, it's the same data science as before, and I think that's what's causing the confusion is, hey, we need data sci…”
Hamel Husain Sep 25, 2025 ▶ 1:19:26
Insight
Husain: General LLM benchmarks do not correlate with product-specific evals
“Up until now, a lot of the big labs understandably focused on general benchmarks, like MMLU score, human eval, things like that, which are very important for foundation models. And, you know, those not very related to product specific evals, like the ones we t…”
Hamel Husain Sep 25, 2025 ▶ 1:21:35
Opinion
Husain: Buying off-the-shelf automated AI eval tools does not work
“The top one is, hey, I can just buy a tool, plug it in, and it'll do the eval for you. Why do I have to worry about this? We live in the age of AI. Can't the AI just eval it? That's the most common misconception. And people want that so much that people do sel…”
Hamel Husain Sep 25, 2025 ▶ 1:24:31
Insight
Husain: Inspecting raw trace data is the highest-ROI activity for AI builders
“Make it as easy as possible, because again, it's the most powerful activity that you can engage in. It's the highest ROI activity you can engage in.”
Hamel Husain Sep 25, 2025 ▶ 1:29:39
Disclosure
Shankar: Establishing initial AI evals takes 3–4 days plus 30 minutes weekly
“Usually I'll spend three to four days really working with whoever to do initial rounds of error analysis, like a lot of labeling, feel like we're in a good place to create the spreadsheet that Hamill had and everyone's kind of on board and convinced, and even …”
Shreya Shankar Sep 25, 2025 ▶ 1:30:46
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.