LLM as a judge
also referred to as: llm-as-a-judge
9 statements across 7 episodes · 4 bullish · 2 bearish · 7 people on the record · first statement Aug 28, 2024 by Nicholas Carlini · across every show →
Everything said about LLM as a judge, oldest first
Aug 28, 2024 positive
Nov 29, 2024 positive
Eugene Yan: LLM-as-a-Judge works reliably when reduced to binary classification
“I think when we simplify it to binary classification metrics, I think it can work. And I think a lot of things can be simplified, like Shreya mentioned, I think a lot of things can be simplified to binary classification metrics. And I've seen evidence of it wo…”
Jan 2, 2025 bullish
Lambert: Open judge models will become core open RL infrastructure
“We already have a bunch of open models that are doing like judge of models and Prometheus and other things that are designed specifically for LM as a judge. And I see that continuing to just become part of this kind of open RL infrastructure.”
Mar 13, 2025 negative
Mar 13, 2025 negative
Husain: General benchmarks for LLM judges provide very little value
“I put very little value in benchmarks, like general benchmarks. It's has some value, but you know, what you really need to do is like measure it in your domain and see if that alum as a judge is more aligned than like an off the shelf LLM. And what I've found …”
Mar 13, 2025
Husain: LLM judges must be validated against domain experts
“One really huge thing about LLM as a judge, people love LLM as a judge, but you really have to make sure that you can trust the LLM as a judge. And so how do you trust and how do you trust anything is that you have to check it and you have to measure how good …”
Jul 5, 2025 neutral
Anthropic finds a single LLM judge outperforms five specialized judges
“They initially started with five LLM as judges. So each one of these points had their own LLM as a judge. They tested the ability and accuracy of that LLM as judge collective to judge, and it actually didn't perform as well as one. So they replaced all of thos…”
Jul 14, 2025 neutral
Dec 31, 2025 bullish
Bissell: Probing internal model features matches LLM-as-a-judge quality at 500x lower cost
“If you ask that model, try to like use it as an LLM as a judge, it's not very good. But if you probe its mind and you sort of detect when the features related to personally identifiable information are firing, that gets you the highest recall of anything. It's…”