Disclosure certainty 3/5 debate potential 2/5

Shankar: Establishing initial AI evals takes 3–4 days plus 30 minutes weekly

Shreya Shankar · Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar · Sep 25, 2025 · at 1:30:46

AI eval instructor Shreya Shankar details the typical upfront and recurring time commitment required to build and maintain evals.

0:00 / 0:52exact quote · 52.8s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“Usually I'll spend three to four days really working with whoever to do initial rounds of error analysis, like a lot of labeling, feel like we're in a good place to create the spreadsheet that Hamill had and everyone's kind of on board and convinced, and even like a few LLM judge evaluators. But this is a one-time cost. Once I figured out how to integrate that in unit tests, or I have like a script that automatically runs it on samples and I will create a cron job to just do this every week. I would say it's like, I don't know. I find myself probably spending more time looking at data because I'm just data hungry like that. I'm so curious. I'm like, I've gained so much from this process and it's like, put me above and beyond in any of my, you know, collaborations with folks. So I want to keep doing it, but I don't have to, I would say like maybe. 30 minutes a week after that.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Shreya Shankar

Insight
Shankar: LLMs fail at initial error analysis due to missing context
“What we usually find when we try to ask an LLM to do this error analysis is it just says the trace looks good because it doesn't have the context needed to understand whether something might be, you know, bad product smell or, you know, not”
Shreya Shankar Sep 25, 2025 ▶ 24:05 Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar
Insight
Shreya Shankar: AI eval rubrics cannot be defined upfront without data
“What's new here is that you can't figure out your rubrics upfront. People's opinions of good and bad change as they review more outputs. They think of failure modes only after seeing 10 outputs they would never have dreamed of in the first place.”
Shreya Shankar Sep 25, 2025 ▶ 1:04:13 Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar
Opinion
Shreya Shankar: AI companies conceal evals because they are competitive moats
“And people don't talk about it because this is their moat, right? So people are not going to go and share all of these things because it makes sense, right? If you are an email writing assistant and you're doing this and you're doing it well, you don't want so…”
Shreya Shankar Sep 25, 2025 ▶ 1:08:23 Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar
Insight
Shankar: LLM judges are reliable when scoped to binary failure modes
“People always think like, oh, this is at least as hard as my problem of creating the original agent, and it's not because you're asking the judge to do one thing, evaluate one failure mode. So the scope of the problem is very small and the output of this LLM j…”
Shreya Shankar Sep 25, 2025 ▶ 50:50 Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar
Insight
Shreya Shankar: AI products typically need only four to seven LLM evals
“For me, like, between four and seven. It's not that many, because a lot of the failure modes, as Hamill said earlier, can be fixed by just fixing your prompt.”
Shreya Shankar Sep 25, 2025 ▶ 1:05:19 Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar
Insight
Shankar: Open code notes must be detailed for LLMs to categorize errors
“This also drives home the point that your open codes have to be detailed, right? You can't just say janky because if the AI is reading janky, it's not going to be able to categorize it. Even a human wouldn't, right? It would have to go and remember why you sai…”
Shreya Shankar Sep 25, 2025 ▶ 42:26 Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.