AI evaluation specialist Shreya Shankar discusses why using LLMs as evaluators is significantly easier and more reliable than creating the original agent.
“People always think like, oh, this is at least as hard as my problem of creating the original agent, and it's not because you're asking the judge to do one thing, evaluate one failure mode. So the scope of the problem is very small and the output of this LLM judge is like pass or fail. So it is a very, very tightly scoped thing that LLM judges are very capable of doing very reliably.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Shreya Shankar
Insight
Shankar: LLMs fail at initial error analysis due to missing context
“What we usually find when we try to ask an LLM to do this error analysis is it just says the trace looks good because it doesn't have the context needed to understand whether something might be, you know, bad product smell or, you know, not”
Shreya ShankarSep 25, 2025▶ 24:05Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar
Insight
Shreya Shankar: AI eval rubrics cannot be defined upfront without data
“What's new here is that you can't figure out your rubrics upfront. People's opinions of good and bad change as they review more outputs. They think of failure modes only after seeing 10 outputs they would never have dreamed of in the first place.”
Shreya ShankarSep 25, 2025▶ 1:04:13Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar
Opinion
Shreya Shankar: AI companies conceal evals because they are competitive moats
“And people don't talk about it because this is their moat, right? So people are not going to go and share all of these things because it makes sense, right? If you are an email writing assistant and you're doing this and you're doing it well, you don't want so…”
Shreya ShankarSep 25, 2025▶ 1:08:23Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar
Insight
Shreya Shankar: AI products typically need only four to seven LLM evals
“For me, like, between four and seven. It's not that many, because a lot of the failure modes, as Hamill said earlier, can be fixed by just fixing your prompt.”
Shreya ShankarSep 25, 2025▶ 1:05:19Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar
Disclosure
Shankar: Establishing initial AI evals takes 3–4 days plus 30 minutes weekly
“Usually I'll spend three to four days really working with whoever to do initial rounds of error analysis, like a lot of labeling, feel like we're in a good place to create the spreadsheet that Hamill had and everyone's kind of on board and convinced, and even …”
Shreya ShankarSep 25, 2025▶ 1:30:46Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar
Insight
Shankar: Open code notes must be detailed for LLMs to categorize errors
“This also drives home the point that your open codes have to be detailed, right? You can't just say janky because if the AI is reading janky, it's not going to be able to categorize it. Even a human wouldn't, right? It would have to go and remember why you sai…”
Shreya ShankarSep 25, 2025▶ 42:26Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 300 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.