LLM Evaluation Prompts

topic on 1 show · 2 statements across 2 episodes

Latent Space

2 statements about LLM Evaluation Prompts, every show

Yan: Optimizing LLM evaluation prompts requires 100 to 400 labeled examples
“I actually think the right number should be maybe a hundred to 400 if you want to be optimizing based on this.”
Eugene Yan Sep 28, 2024 ▶ 34:39 [Paper Club] Who Validates the Validators? Aligning LLM-Judges with Humans (w/ Eugene Yan)
Schulhoff: LLMs Have Number Biases and Require Explicit Rubrics for Evaluation
“These methods are super problematic because there is an incredible amount of instability in them, in the sense that models are biased towards outputting certain numbers, and you generally shouldn't say things like, output your result as a number on a scale of …”
Sander Schulhoff Sep 20, 2024 ▶ 1:02:55 The Ultimate Guide to Prompting - with Sander Schulhoff from LearnPrompting.org

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.