Model Card

topic on 3 shows · 3 statements across 3 episodes

Latent Space Lenny's Podcast Big Technology

3 statements about Model Card, every show

BIG TECHNOLOGY Assertion Supported
Anthropic assesses biological risk by measuring capability uplift for lay users
“So if you look at our model cards, for example, one of the ways we look at risk in the bio domain is comparing the uplift from sort of lay a lay person using the model versus, you know, an expert or a lay person just using the Internet and seeing what the comp…”
Mike Krieger Jun 25, 2026 ▶ 5:12 Anthropic's Labs Lead On Fable's Capabilities + Building AI-Native Products — With Mike Krieger
Nguyen: AI model card benchmark numbers are never apples-to-apples across labs
“None of the numbers are, like, apples to apples. So you actually need to, like, go back to, like, I don't know, like, GPT-E for model card and, like, read the appendix just to, like, make sure that, like, The settings were the same as you're running the settin…”
Karina Nguyen Feb 1, 2025 ▶ 15:11 The Agent Reasoning Interface: Claude, ChatGPT Canvas, Tasks, Operator — with Karina Nguyen, OpenAI
Krieger: AI benchmark evals often contain flawed ground-truth answers
“And some of these model these evals we've seen, like, even the golden answer, I'm like, I'm not sure a human would say it, or, like, I think that math is actually a little wrong, like, getting a hundred percent is gonna be really hard, because even just gradin…”
Mike Krieger Nov 6, 2024 ▶ 16:50 A conversation with OpenAI's CPO Kevin Weil, Anthropic's CPO Mike Krieger, and Sarah Guo

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.