Mar 13, 2025 · 27m · latent-space

[Lightning Pod] Evals: How to Improve AI Consistently — with Hamel Husain and Shreya Shankar

Hamel Husain · 10m spoken Shawn Wang · 6m spoken Shreya Shankar · 5m spoken Alessio Fanelli · 36s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Hamel Husain and Shreya Shankar break down the essentials of AI evaluations, explaining how systematic measurement, grounded synthetic data, domain-calibrated LLM judges, and practical data literacy enable engineers to reliably move AI applications into production.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 30% of the talking time here. How this is scored →

The hosts as informed peer 5.3 Guest teaching 4.0 Guest disagreement 1.6 The hosts pushing back 2.1
05100:0010:0020:001:08–4:13 · The hosts as informed peer 5/10 The Core Motivation Behind Teaching AI Evaluations Swyx introduces the guests, references previous research discussions, and shares his own framework around evals. Shreya outlines why modern AI engineering evaluation differs from traditional MLOps due to data scarcity.4:13–7:06 · The hosts as informed peer 5/10 Course Syllabus and Hands-On Learning Structure Swyx cites presentations from past conferences and navigates the course website. Shreya and Hamel walk through the four-week syllabus, covering synthetic data, LLM judges, and hands-on coding assignments.7:07–9:30 · The hosts as informed peer 4/10 Lessons from Past Courses and Focused Curriculum Design Swyx asks Hamel how this course will differ from his previous high-profile Maven cohort. Hamel explains his desire to avoid a carnival atmosphere, while Shreya playfully pushes back on Hamel being overly humble.9:30–12:04 · The hosts as informed peer 5/10 The Evergreen Nature of Evaluation Principles Alessio asks whether long course lead times risk obsolescence in fast-moving AI. Hamel and Shreya explain that evaluation fundamentals and data literacy principles are evergreen and not tied to temporary API trends.12:04–14:38 · The hosts as informed peer 3/10 Validating LLM Judges and Grounding Synthetic Data Hamel breaks down Shreya's research on LLM judges, arguing there is no free lunch without human expert verification. Shreya explains the fallacy of choosing between slow human annotation and blind LLM data generation.14:39–17:25 · The hosts as informed peer 6/10 Pitfalls of Generic Benchmarks and Dedicated Judge Models Swyx probes into the efficacy of dedicated judge models and scaling judge compute. Hamel dismisses generic public benchmarks, pointing out that 80% of current LLM-as-a-judge implementations fail to add real value without in-domain alignment.17:25–22:56 · The hosts as informed peer 7/10 Custom Annotation Tooling and Practical Data Literacy Hamel describes the high return on investment for building bespoke domain annotation UIs. Swyx demonstrates strong domain awareness by comparing Eugene Yan's Align Eval, Marimo, and Quadratic spreadsheets.22:57–25:23 · The hosts as informed peer 7/10 Institutionalizing Evals and Industry Adoption Frameworks Swyx argues that evals need an Agile Manifesto or Six Sigma style branding movement to achieve widespread institutional buy-in. Hamel questions whether renaming bridges the gap when practitioners do not recognize their core problem.1:08–4:13 · Guest teaching 4/10 The Core Motivation Behind Teaching AI Evaluations Swyx introduces the guests, references previous research discussions, and shares his own framework around evals. Shreya outlines why modern AI engineering evaluation differs from traditional MLOps due to data scarcity.4:13–7:06 · Guest teaching 3/10 Course Syllabus and Hands-On Learning Structure Swyx cites presentations from past conferences and navigates the course website. Shreya and Hamel walk through the four-week syllabus, covering synthetic data, LLM judges, and hands-on coding assignments.7:07–9:30 · Guest teaching 2/10 Lessons from Past Courses and Focused Curriculum Design Swyx asks Hamel how this course will differ from his previous high-profile Maven cohort. Hamel explains his desire to avoid a carnival atmosphere, while Shreya playfully pushes back on Hamel being overly humble.9:30–12:04 · Guest teaching 5/10 The Evergreen Nature of Evaluation Principles Alessio asks whether long course lead times risk obsolescence in fast-moving AI. Hamel and Shreya explain that evaluation fundamentals and data literacy principles are evergreen and not tied to temporary API trends.12:04–14:38 · Guest teaching 6/10 Validating LLM Judges and Grounding Synthetic Data Hamel breaks down Shreya's research on LLM judges, arguing there is no free lunch without human expert verification. Shreya explains the fallacy of choosing between slow human annotation and blind LLM data generation.14:39–17:25 · Guest teaching 5/10 Pitfalls of Generic Benchmarks and Dedicated Judge Models Swyx probes into the efficacy of dedicated judge models and scaling judge compute. Hamel dismisses generic public benchmarks, pointing out that 80% of current LLM-as-a-judge implementations fail to add real value without in-domain alignment.17:25–22:56 · Guest teaching 4/10 Custom Annotation Tooling and Practical Data Literacy Hamel describes the high return on investment for building bespoke domain annotation UIs. Swyx demonstrates strong domain awareness by comparing Eugene Yan's Align Eval, Marimo, and Quadratic spreadsheets.22:57–25:23 · Guest teaching 3/10 Institutionalizing Evals and Industry Adoption Frameworks Swyx argues that evals need an Agile Manifesto or Six Sigma style branding movement to achieve widespread institutional buy-in. Hamel questions whether renaming bridges the gap when practitioners do not recognize their core problem.1:08–4:13 · Guest disagreement 1/10 The Core Motivation Behind Teaching AI Evaluations Swyx introduces the guests, references previous research discussions, and shares his own framework around evals. Shreya outlines why modern AI engineering evaluation differs from traditional MLOps due to data scarcity.4:13–7:06 · Guest disagreement 1/10 Course Syllabus and Hands-On Learning Structure Swyx cites presentations from past conferences and navigates the course website. Shreya and Hamel walk through the four-week syllabus, covering synthetic data, LLM judges, and hands-on coding assignments.7:07–9:30 · Guest disagreement 1/10 Lessons from Past Courses and Focused Curriculum Design Swyx asks Hamel how this course will differ from his previous high-profile Maven cohort. Hamel explains his desire to avoid a carnival atmosphere, while Shreya playfully pushes back on Hamel being overly humble.9:30–12:04 · Guest disagreement 2/10 The Evergreen Nature of Evaluation Principles Alessio asks whether long course lead times risk obsolescence in fast-moving AI. Hamel and Shreya explain that evaluation fundamentals and data literacy principles are evergreen and not tied to temporary API trends.12:04–14:38 · Guest disagreement 2/10 Validating LLM Judges and Grounding Synthetic Data Hamel breaks down Shreya's research on LLM judges, arguing there is no free lunch without human expert verification. Shreya explains the fallacy of choosing between slow human annotation and blind LLM data generation.14:39–17:25 · Guest disagreement 3/10 Pitfalls of Generic Benchmarks and Dedicated Judge Models Swyx probes into the efficacy of dedicated judge models and scaling judge compute. Hamel dismisses generic public benchmarks, pointing out that 80% of current LLM-as-a-judge implementations fail to add real value without in-domain alignment.17:25–22:56 · Guest disagreement 1/10 Custom Annotation Tooling and Practical Data Literacy Hamel describes the high return on investment for building bespoke domain annotation UIs. Swyx demonstrates strong domain awareness by comparing Eugene Yan's Align Eval, Marimo, and Quadratic spreadsheets.22:57–25:23 · Guest disagreement 2/10 Institutionalizing Evals and Industry Adoption Frameworks Swyx argues that evals need an Agile Manifesto or Six Sigma style branding movement to achieve widespread institutional buy-in. Hamel questions whether renaming bridges the gap when practitioners do not recognize their core problem.1:08–4:13 · The hosts pushing back 1/10 The Core Motivation Behind Teaching AI Evaluations Swyx introduces the guests, references previous research discussions, and shares his own framework around evals. Shreya outlines why modern AI engineering evaluation differs from traditional MLOps due to data scarcity.4:13–7:06 · The hosts pushing back 1/10 Course Syllabus and Hands-On Learning Structure Swyx cites presentations from past conferences and navigates the course website. Shreya and Hamel walk through the four-week syllabus, covering synthetic data, LLM judges, and hands-on coding assignments.7:07–9:30 · The hosts pushing back 1/10 Lessons from Past Courses and Focused Curriculum Design Swyx asks Hamel how this course will differ from his previous high-profile Maven cohort. Hamel explains his desire to avoid a carnival atmosphere, while Shreya playfully pushes back on Hamel being overly humble.9:30–12:04 · The hosts pushing back 2/10 The Evergreen Nature of Evaluation Principles Alessio asks whether long course lead times risk obsolescence in fast-moving AI. Hamel and Shreya explain that evaluation fundamentals and data literacy principles are evergreen and not tied to temporary API trends.12:04–14:38 · The hosts pushing back 1/10 Validating LLM Judges and Grounding Synthetic Data Hamel breaks down Shreya's research on LLM judges, arguing there is no free lunch without human expert verification. Shreya explains the fallacy of choosing between slow human annotation and blind LLM data generation.14:39–17:25 · The hosts pushing back 3/10 Pitfalls of Generic Benchmarks and Dedicated Judge Models Swyx probes into the efficacy of dedicated judge models and scaling judge compute. Hamel dismisses generic public benchmarks, pointing out that 80% of current LLM-as-a-judge implementations fail to add real value without in-domain alignment.17:25–22:56 · The hosts pushing back 2/10 Custom Annotation Tooling and Practical Data Literacy Hamel describes the high return on investment for building bespoke domain annotation UIs. Swyx demonstrates strong domain awareness by comparing Eugene Yan's Align Eval, Marimo, and Quadratic spreadsheets.22:57–25:23 · The hosts pushing back 6/10 Institutionalizing Evals and Industry Adoption Frameworks Swyx argues that evals need an Agile Manifesto or Six Sigma style branding movement to achieve widespread institutional buy-in. Hamel questions whether renaming bridges the gap when practitioners do not recognize their core problem.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 42.2% · guest 57.8%0:00 · the hosts 42.2% · guest 57.8%3:00 · the hosts 26.9% · guest 73.1%3:00 · the hosts 26.9% · guest 73.1%6:00 · the hosts 12.4% · guest 87.6%6:00 · the hosts 12.4% · guest 87.6%9:00 · the hosts 16.1% · guest 83.9%9:00 · the hosts 16.1% · guest 83.9%12:00 · the hosts 10.7% · guest 89.3%12:00 · the hosts 10.7% · guest 89.3%15:00 · the hosts 29% · guest 71%15:00 · the hosts 29% · guest 71%18:00 · the hosts 11.1% · guest 88.9%18:00 · the hosts 11.1% · guest 88.9%21:00 · the hosts 53.9% · guest 46.1%21:00 · the hosts 53.9% · guest 46.1%24:00 · the hosts 63.4% · guest 36.6%24:00 · the hosts 63.4% · guest 36.6%27:00 · the hosts 86% · guest 14%27:00 · the hosts 86% · guest 14%
Sharpest disagreement ▶ 16:00 Hamel Rejects Public Benchmarks for LLM Judges

Hamel forcefully dismisses general evaluation benchmarks, asserting that off-the-shelf metrics mean little without testing alignment on domain-specific client data.

Hardest push from the hosts ▶ 24:05 Swyx Doubles Down on Six Sigma Style Certification Framing

When Hamel questions whether branding helps when users lack problem awareness, Swyx pushes back with his thesis that strong naming and manifesto-style movements organically drive widespread industry adoption.

Biggest teaching moment ▶ 12:35 No Free Lunch Without Domain Expert Validation

Hamel and Shreya educate the hosts on why LLM-as-a-judge cannot be trusted blindly, emphasizing that every automated evaluator requires rigorous human baseline validation.

The host holds their own ▶ 20:52 Swyx Details Python-Powered Spreadsheet Tooling

Swyx demonstrates his technical product depth by discussing Quadratic, Marimo, and programmatic spreadsheets as pragmatic minimal viable interfaces for error analysis.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
The Core Motivation Behind Teaching AI Evaluations 5411 Swyx introduces the guests, references previous research discussions, and shares his own framework around evals. Shreya outlines why modern AI engineering evaluation differs from traditional MLOps due to data scarcity.
Course Syllabus and Hands-On Learning Structure 5311 Swyx cites presentations from past conferences and navigates the course website. Shreya and Hamel walk through the four-week syllabus, covering synthetic data, LLM judges, and hands-on coding assignments.
Lessons from Past Courses and Focused Curriculum Design 4211 Swyx asks Hamel how this course will differ from his previous high-profile Maven cohort. Hamel explains his desire to avoid a carnival atmosphere, while Shreya playfully pushes back on Hamel being overly humble.
The Evergreen Nature of Evaluation Principles 5522 Alessio asks whether long course lead times risk obsolescence in fast-moving AI. Hamel and Shreya explain that evaluation fundamentals and data literacy principles are evergreen and not tied to temporary API trends.
Validating LLM Judges and Grounding Synthetic Data 3621 Hamel breaks down Shreya's research on LLM judges, arguing there is no free lunch without human expert verification. Shreya explains the fallacy of choosing between slow human annotation and blind LLM data generation.
Pitfalls of Generic Benchmarks and Dedicated Judge Models 6533 Swyx probes into the efficacy of dedicated judge models and scaling judge compute. Hamel dismisses generic public benchmarks, pointing out that 80% of current LLM-as-a-judge implementations fail to add real value without in-domain alignment.
Custom Annotation Tooling and Practical Data Literacy 7412 Hamel describes the high return on investment for building bespoke domain annotation UIs. Swyx demonstrates strong domain awareness by comparing Eugene Yan's Align Eval, Marimo, and Quadratic spreadsheets.
Institutionalizing Evals and Industry Adoption Frameworks 7326 Swyx argues that evals need an Agile Manifesto or Six Sigma style branding movement to achieve widespread institutional buy-in. Hamel questions whether renaming bridges the gap when practitioners do not recognize their core problem.

Statements from this episode (13)

Insight
Husain: AI builders consistently get stuck moving demos to production
“Anytime that I try to help someone build an AI application, they always get stuck on how to move beyond a demo product. And they get stuck like how to systematically improve things and measure it.”
Hamel Husain Mar 13, 2025 ▶ 1:30
Insight
Shankar: AI evals differ from MLOps due to data scarcity
“The other thing is I think that AI engineering Evaluation or evals here is actually different from MLOps or ML evaluation for traditional ML models. We were in a much more, you know, data rich setting in MLOps. So we were taught to come up with loss metrics or…”
Shreya Shankar Mar 13, 2025 ▶ 3:20
Insight
Husain: AI evaluation principles are evergreen unless AGI arrives
“And I found that, like, the subject is pretty evergreen, because we're not, you know, over the last year and a half, like, the same principles apply. And, you know, we're not really talking about Like, you know, using specific tools and APIs is more of a gener…”
Hamel Husain Mar 13, 2025 ▶ 10:32
Insight
Shankar: LLM failure modes and evaluation techniques have largely stabilized
“I think techniques have stabilized. I think the kinds of failure modes of LLMs, I mean, they're still there, but it's not like changing every single day. We know that LLMs are bad at certain things. We know a little bit more about say limitations of the transf…”
Shreya Shankar Mar 13, 2025 ▶ 11:21
Insight
Husain: LLM judges must be validated against domain experts
“One really huge thing about LLM as a judge, people love LLM as a judge, but you really have to make sure that you can trust the LLM as a judge. And so how do you trust and how do you trust anything is that you have to check it and you have to measure how good …”
Hamel Husain Mar 13, 2025 ▶ 12:34
Insight
Shankar: Grounded synthetic data generation beats slow human annotation
“People don't know how to do anything other than Plan A, which is to go out and try to collect as much real world data as possible and take months because we're going to employ, like, human annotator teams to do this. Or Plan B, which is I'm going to ask an LLM…”
Shreya Shankar Mar 13, 2025 ▶ 13:48
Opinion
Husain: 80% of LLM-as-a-judge implementations are unhelpful
“I feel like a 75% LMS judge because it's low effort is kind of easy, but I would say out of the 75%, 80% is not helpful.”
Hamel Husain Mar 13, 2025 ▶ 14:58
Insight
Husain: General benchmarks for LLM judges provide very little value
“I put very little value in benchmarks, like general benchmarks. It's has some value, but you know, what you really need to do is like measure it in your domain and see if that alum as a judge is more aligned than like an off the shelf LLM. And what I've found …”
Hamel Husain Mar 13, 2025 ▶ 16:05
Insight
Husain: Custom annotation web apps yield massive ROI for AI evals
“One counterintuitive thing that has an extreme value That people kind of discover maybe accidentally are, you know, if they're working with us, they discover very fast is Is this, there's a really, so you really want to look at your data a lot, and there's a r…”
Hamel Husain Mar 13, 2025 ▶ 17:31
Insight
Husain: Basic spreadsheet skills are enough to run AI evals
“The foundation of evals is error analysis. So like looking at your data and doing data analysis on your traces. So a lot of people, when we say data literacy, that can mean, that can sound scary, but it can come from a lot of different places. It can be, you c…”
Hamel Husain Mar 13, 2025 ▶ 21:37
Insight
Husain: Teams rarely associate AI underperformance with a lack of evals
“Cause like one thing that I wrestle with is like evals is a solution, but the problem is, okay, your AI doesn't work, or it doesn't work as well as you want it to. And people don't associate the solution with the problem cleanly enough. Cause they don't know. …”
Hamel Husain Mar 13, 2025 ▶ 23:49
Prediction Not checkable as stated
Husain: HCI and workflow evaluations will enter AI tools by 2027
“If I were to fast forward one or two years, I would expect to see those and all the tools. It just hasn't arrived yet.”
Hamel Husain Mar 13, 2025 ▶ 26:15
Insight
Shankar: Evals are necessary to train AI reasoning models
“You need evals to train your reasoning models.”
Shreya Shankar Mar 13, 2025 ▶ 27:13
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.