Performance Evals
topic on 1 show · 1 statements across 1 episodes
1 statements about Performance Evals, every show
Clark: High-level LLM evaluations mask undesired AI system behaviors
“We're seeing people do the exact same thing again today with LLMs, where they're focusing on these high-level metrics, these end outputs, these performance evals, and that ends up masking all of these potentially undesired behaviors within the system itself.”