Model Card
topic on 3 shows · 3 statements across 3 episodes
Latent Space
Lenny's Podcast
Big Technology
3 statements about Model Card, every show
Anthropic assesses biological risk by measuring capability uplift for lay users
“So if you look at our model cards, for example, one of the ways we look at risk in the bio domain is comparing the uplift from sort of lay a lay person using the model versus, you know, an expert or a lay person just using the Internet and seeing what the comp…”
Nguyen: AI model card benchmark numbers are never apples-to-apples across labs
“None of the numbers are, like, apples to apples. So you actually need to, like, go back to, like, I don't know, like, GPT-E for model card and, like, read the appendix just to, like, make sure that, like, The settings were the same as you're running the settin…”
Krieger: AI benchmark evals often contain flawed ground-truth answers
“And some of these model these evals we've seen, like, even the golden answer, I'm like, I'm not sure a human would say it, or, like, I think that math is actually a little wrong, like, getting a hundred percent is gonna be really hard, because even just gradin…”