AI consultant Hamel Husain discusses evaluation methodology and common pitfalls when deploying LLMs as automated judges.
Opinion
Husain: 80% of LLM-as-a-judge implementations are unhelpful
“I feel like a 75% LMS judge because it's low effort is kind of easy, but I would say out of the 75%, 80% is not helpful.”
Insight
Husain: General benchmarks for LLM judges provide very little value
“I put very little value in benchmarks, like general benchmarks. It's has some value, but you know, what you really need to do is like measure it in your domain and see if that alum as a judge is more aligned than like an off the shelf LLM. And what I've found …”
Insight
Husain: AI builders consistently get stuck moving demos to production
“Anytime that I try to help someone build an AI application, they always get stuck on how to move beyond a demo product. And they get stuck like how to systematically improve things and measure it.”
Insight
Husain: AI evaluation principles are evergreen unless AGI arrives
“And I found that, like, the subject is pretty evergreen, because we're not, you know, over the last year and a half, like, the same principles apply. And, you know, we're not really talking about Like, you know, using specific tools and APIs is more of a gener…”
Insight
Husain: Custom annotation web apps yield massive ROI for AI evals
“One counterintuitive thing that has an extreme value That people kind of discover maybe accidentally are, you know, if they're working with us, they discover very fast is Is this, there's a really, so you really want to look at your data a lot, and there's a r…”
Insight
Husain: Basic spreadsheet skills are enough to run AI evals
“The foundation of evals is error analysis. So like looking at your data and doing data analysis on your traces. So a lot of people, when we say data literacy, that can mean, that can sound scary, but it can come from a lot of different places. It can be, you c…”