LLM evaluation prompts
2 statements across 2 episodes · 0 bullish · 1 bearish · 2 people on the record · first statement Sep 20, 2024 by Sander Schulhoff · across every show →
Everything said about LLM evaluation prompts, oldest first
Sep 20, 2024 negative
Schulhoff: LLMs Have Number Biases and Require Explicit Rubrics for Evaluation
“These methods are super problematic because there is an incredible amount of instability in them, in the sense that models are biased towards outputting certain numbers, and you generally shouldn't say things like, output your result as a number on a scale of …”