LLM as a judge

also referred to as: llm-as-a-judge

9 statements across 7 episodes · 4 bullish · 2 bearish · 7 people on the record · first statement Aug 28, 2024 by Nicholas Carlini · across every show →

Everything said about LLM as a judge, oldest first

Aug 28, 2024 positive
Insight
Carlini: LLM-as-a-judge is almost always accurate when prompted correctly
“I've inspected the outputs of these and like, they're almost always correct. If you sort of, if you ask the model to judge these things in the right way, they're very good at being able to tell this.”
Nicholas Carlini Aug 28, 2024 ▶ 42:52 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Nov 29, 2024 positive
Insight
Eugene Yan: LLM-as-a-Judge works reliably when reduced to binary classification
“I think when we simplify it to binary classification metrics, I think it can work. And I think a lot of things can be simplified, like Shreya mentioned, I think a lot of things can be simplified to binary classification metrics. And I've seen evidence of it wo…”
Eugene Yan Nov 29, 2024 ▶ 41:04 [Paper Club] DocETL: Agentic Query Rewriting + Eval for Complex Document Processing w Shreya Shankar
Jan 2, 2025 bullish
Prediction Not checkable as stated
Lambert: Open judge models will become core open RL infrastructure
“We already have a bunch of open models that are doing like judge of models and Prometheus and other things that are designed specifically for LM as a judge. And I see that continuing to just become part of this kind of open RL infrastructure.”
Nathan Lambert Jan 2, 2025 ▶ 13:25 The State of Reasoning — from Nathan Lambert, Interconnects/AI2 [LS Live @ NeurIPS 2024]
Mar 13, 2025 negative
Opinion
Husain: 80% of LLM-as-a-judge implementations are unhelpful
“I feel like a 75% LMS judge because it's low effort is kind of easy, but I would say out of the 75%, 80% is not helpful.”
Hamel Husain Mar 13, 2025 ▶ 14:58 [Lightning Pod] Evals: How to Improve AI Consistently — with Hamel Husain and Shreya Shankar
Mar 13, 2025 negative
Insight
Husain: General benchmarks for LLM judges provide very little value
“I put very little value in benchmarks, like general benchmarks. It's has some value, but you know, what you really need to do is like measure it in your domain and see if that alum as a judge is more aligned than like an off the shelf LLM. And what I've found …”
Hamel Husain Mar 13, 2025 ▶ 16:05 [Lightning Pod] Evals: How to Improve AI Consistently — with Hamel Husain and Shreya Shankar
Mar 13, 2025
Insight
Husain: LLM judges must be validated against domain experts
“One really huge thing about LLM as a judge, people love LLM as a judge, but you really have to make sure that you can trust the LLM as a judge. And so how do you trust and how do you trust anything is that you have to check it and you have to measure how good …”
Hamel Husain Mar 13, 2025 ▶ 12:34 [Lightning Pod] Evals: How to Improve AI Consistently — with Hamel Husain and Shreya Shankar
Jul 5, 2025 neutral
Assertion Supported
Anthropic finds a single LLM judge outperforms five specialized judges
“They initially started with five LLM as judges. So each one of these points had their own LLM as a judge. They tested the ability and accuracy of that LLM as judge collective to judge, and it actually didn't perform as well as one. So they replaced all of thos…”
Dylan Davis Jul 5, 2025 ▶ 11:03 ⚡️Anthropic vs Cognition on Multi-Agents: A Breakdown with Dylan Davis
Jul 14, 2025 neutral
Insight
Changing LLM judge models can heavily alter benchmark leaderboards
“The leaderboard, if you are using different judges with different models, it can, there can be heavy shakeup of the leaderboard also.”
Pratik Bhavsar Jul 14, 2025 ▶ 22:45 ⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
Dec 31, 2025 bullish
Insight
Bissell: Probing internal model features matches LLM-as-a-judge quality at 500x lower cost
“If you ask that model, try to like use it as an LLM as a judge, it's not very good. But if you probe its mind and you sort of detect when the features related to personally identifiable information are firing, that gets you the highest recall of anything. It's…”
Mark Bissell Dec 31, 2025 ▶ 9:53 [State of MechInterp] SAEs in Production, Circuit Tracing, AI4Science, "Pragmatic" Interp — Goodfire
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.