Sep 28, 2024 · 1h 0m · latent-space
[Paper Club] Who Validates the Validators? Aligning LLM-Judges with Humans (w/ Eugene Yan)
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
Eugene Yan leads a Paper Club discussion analyzing the paper 'Who Validates the Validators? Aligning LLM-Judges with Humans' by Shreya Shankar et al., exploring methodologies for constructing, aligning, and benchmarking automated LLM evaluators. The session demonstrates how human-in-the-loop inspection, decomposed assertions, and interactive tooling like the prototype 'LABEL' app reliably optimize evaluator performance.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
SPEAKER_12 offers mild pushback against Eugene's binary classification premise, arguing that even objective tasks like medical instruction following involve subjective taste among domain experts.
Hardest push from the hosts ▶ 59:34 Host pushes past Gorilla paper to newer Berkeley benchmarksHost Alessio Fanelli steers the discussion by asserting that the original Gorilla paper is no longer as relevant compared to the Berkeley Function Calling Leaderboard v2 and v3.
Biggest teaching moment ▶ 8:54 Translating paper metrics into standard ML terminologyEugene clarifies the paper's idiosyncratic terminology by translating coverage directly to recall and false failure rate to one minus precision.
The host holds their own ▶ 59:03 Host outlines Berkeley LLM ecosystem contextAlessio Fanelli takes over the wrap-up, demonstrating broad industry context across Berkeley function calling leaderboards, LMSYS, and upcoming research pipelines.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Welcome, Session Logistics, and Agenda Overview | 0 | 1 | 0 | 0 | The episode begins with housekeeping logistics by the host before Eugene Yan introduces the 'Who Validates the Validators?' paper and explains the core chicken-and-egg problem of defining evaluation criteria without inspecting model outputs. | |
| EvalGen Architecture, Assertion Methods, and Benchmark Results | 0 | 0 | 0 | 0 | Eugene Yan monologues through the architecture of EvalGen, breaking down metrics such as coverage (recall) and false failure rate (one minus precision), and reviewing benchmark results across medical and product pipelines. | |
| User Study Insights and Practitioner Reflections on Alignment | 0 | 0 | 0 | 0 | Eugene reviews the paper's user study involving nine industry practitioners, discussing the iterative nature of criteria definition, human labeling bias, and skepticism toward using LLMs as production evaluators. | |
| Demonstration of the LABEL Prototype Application | 0 | 0 | 0 | 0 | Eugene conducts a live demo of his prototype application, LABEL, walking through labeling, evaluation, and hyper-prompt optimization modes while discussing the technical stack with attendees. | |
| Technical Discussion: Binary Judgments, Re-Rankers, and Task Types | 0 | 1 | 1 | 0 | Attendees ask technical questions regarding binary distillations, re-rankers versus classifiers, and the boundary between objective metrics and subjective taste in LLM evaluations. |