Evaluation Metrics
topic on 3 shows · 4 statements across 4 episodes
Latent Space
Lenny's Podcast
the MAD Podcast
4 statements about Evaluation Metrics, every show
Reganti: Predefined AI evals only catch anticipated failure patterns
“The issue with just building a bunch of evaluation metrics and then having them in production is evaluation metrics catch only the errors that you're already aware of, but there can be a lot of emerging patterns that you understand. Only after you put things i…”
Chip Huyen: AI evaluation metrics must be derived backward from business use cases
“For applications it's really, really important to understand the use cases well, so they can design like the set of metrics and then you can walk backward from that and map it to like the model metrics.”
Albrecht: LLM emergence is an artifact of non-linear evaluation metrics
“This emergent behavior that you're seeing, Is not really emergent behavior, but is really a function of the evaluation metrics that we're using.”
Kanjun Qiu: AI emergent capabilities are artifacts of evaluation metric design
“If your metric is smooth you actually see slow performance improvement over time, and if your metric is relatively discrete or not smooth, that's where you see the emergence, and it's actually more about the evaluation metric Than about the emergence of the ca…”