LLM judge

also referred to as: llm judges

3 statements across 2 episodes · 0 bullish · 2 bearish · 3 people on the record · first statement Sep 25, 2025 by Hamel Husain · across every show →

Everything said about LLM judge, oldest first

Sep 25, 2025 negative
Insight
Husain: Raw human-judge agreement is a misleading metric for AI evals
“Now, one thing you should know as a product manager is a lot of people go straight to this, like, agreement. They say, okay, my judge agrees with the human at some percentage of the time. Now that sounds appealing, but it's a very dangerous metric to use becau…”
Hamel Husain Sep 25, 2025 ▶ 58:24 Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar
Sep 25, 2025 neutral
Insight
Shreya Shankar: AI products typically need only four to seven LLM evals
“For me, like, between four and seven. It's not that many, because a lot of the failure modes, as Hamill said earlier, can be fixed by just fixing your prompt.”
Shreya Shankar Sep 25, 2025 ▶ 1:05:19 Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar
Jan 11, 2026 negative
Insight
Reganti: LLM judges fail in complex AI cases due to emerging patterns
“When you go to complex use cases, it's incredibly hard to build LLM judges because you see a lot of emerging patterns. If you build a judge that would you know, test for verbosity or something like that, it turns out that you're seeing newer patterns that your…”
Aishwarya Reganti (Ash) Jan 11, 2026 ▶ 40:07 Why most AI products fail: Lessons from 50+ AI deployments at OpenAI, Google & Amazon
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.