AI Evals

topic on 5 shows · 8 statements across 6 episodes

the Y Combinator Startup Podcast Latent Space Lenny's Podcast No Priors the a16z Podcast

8 statements about AI Evals, every show

a16z Insight
Building evals is the hardest part of model routing
“Really the hardest part of routing is building the evals and trying to determine in what places a set of intelligences should be used for a particular application.”
Ryan Chi Sep 9, 2026 ▶ 16:04 Inside the Race to Measure Frontier Intelligence
Husain: Raw human-judge agreement is a misleading metric for AI evals
“Now, one thing you should know as a product manager is a lot of people go straight to this, like, agreement. They say, okay, my judge agrees with the human at some percentage of the time. Now that sounds appealing, but it's a very dangerous metric to use becau…”
Hamel Husain Sep 25, 2025 ▶ 58:24 Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar
Husain: Buying off-the-shelf automated AI eval tools does not work
“The top one is, hey, I can just buy a tool, plug it in, and it'll do the eval for you. Why do I have to worry about this? We live in the age of AI. Can't the AI just eval it? That's the most common misconception. And people want that so much that people do sel…”
Hamel Husain Sep 25, 2025 ▶ 1:24:31 Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar
Shankar: Establishing initial AI evals takes 3–4 days plus 30 minutes weekly
“Usually I'll spend three to four days really working with whoever to do initial rounds of error analysis, like a lot of labeling, feel like we're in a good place to create the spreadsheet that Hamill had and everyone's kind of on board and convinced, and even …”
Shreya Shankar Sep 25, 2025 ▶ 1:30:46 Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar
Field: Designers and PMs, not just researchers, should build AI evals
“As you're, you know, doing developing a model or you're developing research ideas, You have to have good evals, and usually the researchers are the ones building those, and I think that's kind of just the wrong model. For us, at least, designers, my point of v…”
Dylan Field Aug 8, 2025 ▶ 28:00 Dylan Field: Scaling Figma and the Future of Design · Y Combinator
Alberti: Small numbers of manually curated evals beat large eval sets
“I think usually like small numbers of very high quality evals are the way to go and like we curate them manually and make sure they're really good.”
Silas Alberti May 21, 2025 ▶ 17:39 DeepWiki: The GitHub Encyclopedia
NO PRIORS Prediction Not checkable as stated
Foody: Software engineer evals will take years to build
“All the things that go into making a good software engineer, that's gonna be really hard to do. Like, I think it's going to be a years long build out for even some of the verifiable domains. Cause there's so much that goes into a good software engineer of like…”
Brendan Foody Apr 10, 2025 ▶ 16:14 No Priors Ep. 110 | With Mercor CEO and Co-Founder Brendan Foody
Husain: Basic spreadsheet skills are enough to run AI evals
“The foundation of evals is error analysis. So like looking at your data and doing data analysis on your traces. So a lot of people, when we say data literacy, that can mean, that can sound scary, but it can come from a lot of different places. It can be, you c…”
Hamel Husain Mar 13, 2025 ▶ 21:37 [Lightning Pod] Evals: How to Improve AI Consistently — with Hamel Husain and Shreya Shankar

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.