Sep 28, 2024 · 1h 0m · latent-space

[Paper Club] Who Validates the Validators? Aligning LLM-Judges with Humans (w/ Eugene Yan)

Eugene Yan · 40m spoken Shreya Shankar · 1m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Eugene Yan leads a Paper Club discussion analyzing the paper 'Who Validates the Validators? Aligning LLM-Judges with Humans' by Shreya Shankar et al., exploring methodologies for constructing, aligning, and benchmarking automated LLM evaluators. The session demonstrates how human-in-the-loop inspection, decomposed assertions, and interactive tooling like the prototype 'LABEL' app reliably optimize evaluator performance.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 0.0 Guest teaching 0.4 Guest disagreement 0.2 The hosts pushing back 0.0
05100:0015:0030:0045:001:00:000:00–7:03 · The hosts as informed peer 0/10 Welcome, Session Logistics, and Agenda Overview The episode begins with housekeeping logistics by the host before Eugene Yan introduces the 'Who Validates the Validators?' paper and explains the core chicken-and-egg problem of defining evaluation criteria without inspecting model outputs.7:03–15:44 · The hosts as informed peer 0/10 EvalGen Architecture, Assertion Methods, and Benchmark Results Eugene Yan monologues through the architecture of EvalGen, breaking down metrics such as coverage (recall) and false failure rate (one minus precision), and reviewing benchmark results across medical and product pipelines.15:44–30:08 · The hosts as informed peer 0/10 User Study Insights and Practitioner Reflections on Alignment Eugene reviews the paper's user study involving nine industry practitioners, discussing the iterative nature of criteria definition, human labeling bias, and skepticism toward using LLMs as production evaluators.30:08–41:56 · The hosts as informed peer 0/10 Demonstration of the LABEL Prototype Application Eugene conducts a live demo of his prototype application, LABEL, walking through labeling, evaluation, and hyper-prompt optimization modes while discussing the technical stack with attendees.41:56–55:07 · The hosts as informed peer 0/10 Technical Discussion: Binary Judgments, Re-Rankers, and Task Types Attendees ask technical questions regarding binary distillations, re-rankers versus classifiers, and the boundary between objective metrics and subjective taste in LLM evaluations.0:00–7:03 · Guest teaching 1/10 Welcome, Session Logistics, and Agenda Overview The episode begins with housekeeping logistics by the host before Eugene Yan introduces the 'Who Validates the Validators?' paper and explains the core chicken-and-egg problem of defining evaluation criteria without inspecting model outputs.7:03–15:44 · Guest teaching 0/10 EvalGen Architecture, Assertion Methods, and Benchmark Results Eugene Yan monologues through the architecture of EvalGen, breaking down metrics such as coverage (recall) and false failure rate (one minus precision), and reviewing benchmark results across medical and product pipelines.15:44–30:08 · Guest teaching 0/10 User Study Insights and Practitioner Reflections on Alignment Eugene reviews the paper's user study involving nine industry practitioners, discussing the iterative nature of criteria definition, human labeling bias, and skepticism toward using LLMs as production evaluators.30:08–41:56 · Guest teaching 0/10 Demonstration of the LABEL Prototype Application Eugene conducts a live demo of his prototype application, LABEL, walking through labeling, evaluation, and hyper-prompt optimization modes while discussing the technical stack with attendees.41:56–55:07 · Guest teaching 1/10 Technical Discussion: Binary Judgments, Re-Rankers, and Task Types Attendees ask technical questions regarding binary distillations, re-rankers versus classifiers, and the boundary between objective metrics and subjective taste in LLM evaluations.0:00–7:03 · Guest disagreement 0/10 Welcome, Session Logistics, and Agenda Overview The episode begins with housekeeping logistics by the host before Eugene Yan introduces the 'Who Validates the Validators?' paper and explains the core chicken-and-egg problem of defining evaluation criteria without inspecting model outputs.7:03–15:44 · Guest disagreement 0/10 EvalGen Architecture, Assertion Methods, and Benchmark Results Eugene Yan monologues through the architecture of EvalGen, breaking down metrics such as coverage (recall) and false failure rate (one minus precision), and reviewing benchmark results across medical and product pipelines.15:44–30:08 · Guest disagreement 0/10 User Study Insights and Practitioner Reflections on Alignment Eugene reviews the paper's user study involving nine industry practitioners, discussing the iterative nature of criteria definition, human labeling bias, and skepticism toward using LLMs as production evaluators.30:08–41:56 · Guest disagreement 0/10 Demonstration of the LABEL Prototype Application Eugene conducts a live demo of his prototype application, LABEL, walking through labeling, evaluation, and hyper-prompt optimization modes while discussing the technical stack with attendees.41:56–55:07 · Guest disagreement 1/10 Technical Discussion: Binary Judgments, Re-Rankers, and Task Types Attendees ask technical questions regarding binary distillations, re-rankers versus classifiers, and the boundary between objective metrics and subjective taste in LLM evaluations.0:00–7:03 · The hosts pushing back 0/10 Welcome, Session Logistics, and Agenda Overview The episode begins with housekeeping logistics by the host before Eugene Yan introduces the 'Who Validates the Validators?' paper and explains the core chicken-and-egg problem of defining evaluation criteria without inspecting model outputs.7:03–15:44 · The hosts pushing back 0/10 EvalGen Architecture, Assertion Methods, and Benchmark Results Eugene Yan monologues through the architecture of EvalGen, breaking down metrics such as coverage (recall) and false failure rate (one minus precision), and reviewing benchmark results across medical and product pipelines.15:44–30:08 · The hosts pushing back 0/10 User Study Insights and Practitioner Reflections on Alignment Eugene reviews the paper's user study involving nine industry practitioners, discussing the iterative nature of criteria definition, human labeling bias, and skepticism toward using LLMs as production evaluators.30:08–41:56 · The hosts pushing back 0/10 Demonstration of the LABEL Prototype Application Eugene conducts a live demo of his prototype application, LABEL, walking through labeling, evaluation, and hyper-prompt optimization modes while discussing the technical stack with attendees.41:56–55:07 · The hosts pushing back 0/10 Technical Discussion: Binary Judgments, Re-Rankers, and Task Types Attendees ask technical questions regarding binary distillations, re-rankers versus classifiers, and the boundary between objective metrics and subjective taste in LLM evaluations.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 52:39 SPEAKER_12 challenges objective vs subjective separation

SPEAKER_12 offers mild pushback against Eugene's binary classification premise, arguing that even objective tasks like medical instruction following involve subjective taste among domain experts.

Hardest push from the hosts ▶ 59:34 Host pushes past Gorilla paper to newer Berkeley benchmarks

Host Alessio Fanelli steers the discussion by asserting that the original Gorilla paper is no longer as relevant compared to the Berkeley Function Calling Leaderboard v2 and v3.

Biggest teaching moment ▶ 8:54 Translating paper metrics into standard ML terminology

Eugene clarifies the paper's idiosyncratic terminology by translating coverage directly to recall and false failure rate to one minus precision.

The host holds their own ▶ 59:03 Host outlines Berkeley LLM ecosystem context

Alessio Fanelli takes over the wrap-up, demonstrating broad industry context across Berkeley function calling leaderboards, LMSYS, and upcoming research pipelines.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Welcome, Session Logistics, and Agenda Overview 0100 The episode begins with housekeeping logistics by the host before Eugene Yan introduces the 'Who Validates the Validators?' paper and explains the core chicken-and-egg problem of defining evaluation criteria without inspecting model outputs.
EvalGen Architecture, Assertion Methods, and Benchmark Results 0000 Eugene Yan monologues through the architecture of EvalGen, breaking down metrics such as coverage (recall) and false failure rate (one minus precision), and reviewing benchmark results across medical and product pipelines.
User Study Insights and Practitioner Reflections on Alignment 0000 Eugene reviews the paper's user study involving nine industry practitioners, discussing the iterative nature of criteria definition, human labeling bias, and skepticism toward using LLMs as production evaluators.
Demonstration of the LABEL Prototype Application 0000 Eugene conducts a live demo of his prototype application, LABEL, walking through labeling, evaluation, and hyper-prompt optimization modes while discussing the technical stack with attendees.
Technical Discussion: Binary Judgments, Re-Rankers, and Task Types 0110 Attendees ask technical questions regarding binary distillations, re-rankers versus classifiers, and the boundary between objective metrics and subjective taste in LLM evaluations.

Statements from this episode (8)

Insight
Yan: LLM evaluation criteria cannot be determined without inspecting real outputs
“What they're saying is that it is impossible to completely determine good evaluation criteria without actually looking at LLM outputs. So essentially the point is you have to look at the data before you overcome evaluation criteria.”
Eugene Yan Sep 28, 2024 ▶ 4:01
Assertion Supported
Yan: EvalGen requires fewer assertions than SPADE for comparable evaluation
“Evalgen only needed three assertions, be quote assertions, or LLM prompts. So essentially, it needed less than spade, which needed five to get comparable results. Now, when we look at the product pipeline we see that evalgen only requires four assertions, whic…”
Eugene Yan Sep 28, 2024 ▶ 13:35
Insight
Yan: Annotators must update past grades when evaluation criteria drift
“If you find that your criteria has drifted, instead of trying to maintain the same criteria, And aligning to the previous grades. Instead, what we should do is we should revisit those previous grades and fix it because it's an iterative process.”
Eugene Yan Sep 28, 2024 ▶ 25:27
Opinion
Yan: Optimizing LLM evaluation prompts requires 100 to 400 labeled examples
“I actually think the right number should be maybe a hundred to 400 if you want to be optimizing based on this.”
Eugene Yan Sep 28, 2024 ▶ 34:39
Opinion
Yan: Using LLMs as evaluators is the only way to scale
“I know that we have to use an LLM as an evaluator. There's no way around it. If we want to scale, I think that's the only way.”
Eugene Yan Sep 28, 2024 ▶ 44:00
Insight
Yan: Pairwise evaluation fails for objective metrics like factuality
“The reason why pairwise preferences cannot work is that if you give two things that are both factual or if you give two things that are both non-factual, you would say that one is better than the other, but it still doesn't meet the bar of being factual enough…”
Eugene Yan Sep 28, 2024 ▶ 49:36
Disclosure
Shankar: New LLM evaluation framework to be open-sourced in ChainForge
“Not yet. It's not out yet, but we will. So the conference is in three weeks. We have to have it out by then, but it'll be implemented in chainforge.ai, which is an open source LLM pipeline building tool.”
Shreya Shankar Sep 28, 2024 ▶ 56:49
Insight
Shankar: Decompose LLM pipelines into unit tasks with standalone intermediate assertions
“The idea is to have each node in your graph kind of be a standalone, do a standalone thing that you can have standalone assertions for. And if you think about, you know, infinitely many inputs flowing through your pipeline, there's going to be some fraction of…”
Shreya Shankar Sep 28, 2024 ▶ 57:35
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.