LLM judge
also referred to as: llm judges
3 statements across 2 episodes · 0 bullish · 2 bearish · 3 people on the record · first statement Sep 25, 2025 by Hamel Husain · across every show →
Everything said about LLM judge, oldest first
Sep 25, 2025 negative
Husain: Raw human-judge agreement is a misleading metric for AI evals
“Now, one thing you should know as a product manager is a lot of people go straight to this, like, agreement. They say, okay, my judge agrees with the human at some percentage of the time. Now that sounds appealing, but it's a very dangerous metric to use becau…”
Sep 25, 2025 neutral
Jan 11, 2026 negative
Reganti: LLM judges fail in complex AI cases due to emerging patterns
“When you go to complex use cases, it's incredibly hard to build LLM judges because you see a lot of emerging patterns. If you build a judge that would you know, test for verbosity or something like that, it turns out that you're seeing newer patterns that your…”