LLM Evaluation Prompts
topic on 1 show · 2 statements across 2 episodes
2 statements about LLM Evaluation Prompts, every show
Yan: Optimizing LLM evaluation prompts requires 100 to 400 labeled examples
“I actually think the right number should be maybe a hundred to 400 if you want to be optimizing based on this.”
Schulhoff: LLMs Have Number Biases and Require Explicit Rubrics for Evaluation
“These methods are super problematic because there is an incredible amount of instability in them, in the sense that models are biased towards outputting certain numbers, and you generally shouldn't say things like, output your result as a number on a scale of …”