Sep 25, 2025 · 1h 46m · lennys-podcast
Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this masterclass, AI evaluation experts Hamel Husain and Shreya Shankar demonstrate how product builders can transition from guesswork and 'vibe checks' to systematic, high-ROI error analysis and automated evaluators.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Lenny holds 21.7% of the talking time here. How this is scored →
speaking balance: gold is Lenny, purple is the guest (3 minute bins)
Shreya and Hamel push back strongly against the viral narrative that top products like Claude Code operate solely on vibe checks without rigorous foundational evals and telemetry.
Hardest push from Lenny ▶ 8:29 Lenny Testing the Unit Test MetaphorLenny presses on whether evals can simply be thought of as standard software unit tests, challenging the guests to clarify the exact functional distinction.
Biggest teaching moment ▶ 58:20 Exposing the Trap of High Percentage Agreement in EvalsHamel demonstrates why a product manager celebrating a 90% judge agreement metric might actually be masking massive failure rates on rare edge cases without a confusion matrix.
Lenny holds their own ▶ 1:00:56 Connecting Automated Judges Directly to Product Requirements DocumentsLenny synthesizes the guests' complex prompt engineering mechanics into the core product management discipline, demonstrating how judge prompts function as continuously executing dynamic PRDs.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Lenny as informed peer | Guest teaching | Guest disagreement | Lenny pushing back | Why |
|---|---|---|---|---|---|---|
| Preview: Why AI Evals Are the Hottest Skill | 3 | 1 | 0 | 0 | Lenny delivers an introductory monologue summarizing the rising importance of AI evals and framing Hamel and Shreya's course on Maven before cutting to sponsor reads. | |
| Sponsor: Fin AI Agent by Intercom | 4 | 5 | 1 | 1 | Lenny proposes that evals are simply unit tests for AI, prompting Shreya to gently reframe the concept into a much wider spectrum of data analytics and qualitative discovery. | |
| Demonstrating Error Analysis and Open Coding on Real Traces | 3 | 6 | 1 | 0 | Hamel leads a detailed walkthrough of real-world conversational traces from Nurture Boss, instructing Lenny on open coding and capturing upstream error notes manually. | |
| The Limits of Automation and the Benevolent Dictator | 3 | 6 | 2 | 1 | Shreya warns against early LLM automation for note-taking due to lack of domain context, and Hamel explains the necessity of appointing a single benevolent dictator over a committee. | |
| Synthesizing Open Codes into Axial Categories | 4 | 5 | 1 | 0 | The guests explain theoretical saturation and demonstrate clustering open codes into axial categories with LLMs, referencing foundational machine learning error analysis roots from Andrew Ng. | |
| Quantifying Error Distributions and Prioritizing Fixes | 3 | 5 | 0 | 0 | Hamel demonstrates aggregating categorized errors using pivot tables to quantitatively prioritize application fixes before jumping into automated testing. | |
| Code-Based Evaluators Versus LLM-as-a-Judge | 4 | 6 | 2 | 0 | Shreya and Hamel distinguish deterministic code evaluators from LLM judges, strongly criticizing 1-to-5 Likert scales in favor of strict binary pass-fail metrics. | |
| Sponsor: Mercury Financial Services for Startups | 4 | 6 | 1 | 0 | After an ad break, Hamel details aligning LLM judges with human evaluations and warns why overall percentage agreement is misleading without an error confusion matrix. | |
| Evals as Dynamic PRDs and the Impact of Criteria Drift | 6 | 5 | 1 | 1 | Lenny makes an insightful connection between LLM judge prompts and dynamic PRDs, leading Shreya to discuss her research on criteria drift during model evaluation. | |
| Debating Vibe Checks, Coding Agents, and A/B Testing | 5 | 6 | 3 | 2 | The guests dismantle online debates pitting vibes and A/B testing against evals, explaining that coding agents benefit from unique developer-dogfooding dynamics that do not generalize. | |
| Common Misconceptions, Custom Tools, and Practical Advice | 4 | 5 | 1 | 0 | Hamel and Shreya emphasize looking directly at raw traces and show how simple custom internal interfaces maximize iteration velocity and product ROI. | |
| Overview of the Maven AI Evals Course and Resources | 4 | 4 | 0 | 0 | The guests outline their Maven curriculum, supplemental comprehensive handbook, and custom AI assistant bot trained on their entire corpus. | |
| Lightning Round: Books, Culture, Tools, and Mottos | 3 | 3 | 0 | 0 | The interview concludes with a lighthearted lightning round touching on textbooks, fiction, TV shows, favorite developer tools, and mutual admiration between the guests. |