Evaluation

topic on 6 shows · 8 statements across 8 episodes

the Y Combinator Startup Podcast Latent Space Lenny's Podcast the MAD Podcast the a16z Podcast 20VC

8 statements about Evaluation, every show

a16z Insight
Complex AI evals require smaller sample sizes and broader criteria
“Evaluations as they become more complex, Have a fewer sample size, but a larger set of criteria or expectations of them.”
Ryan Chi Sep 9, 2026 ▶ 14:40 Inside the Race to Measure Frontier Intelligence
20VC Insight
Joelle Pineau: AI benchmarks should be treated as narrow unit tests
“Think of evaluations as, like, unit test for the performance of your system. I mean, software engineers will know what that is, right? Like, you run through that evaluation, and that gives you, like, a signal of how the system is doing in a particular dimensio…”
Joelle Pineau Nov 3, 2025 ▶ 47:42 Cohere's Chief AI Officer, Joelle Pineau: Why Scaling Laws Will Continue & Future of Synthetic Data · 20VC with Harry Stebbings
Y COMBINATOR Disclosure
Heller: Casetext derived most evals from customer failures, not lab tests
“We've added much more evals at this point from real things that happened to real customers than the ones we came up with in the lab.”
Jake Heller Oct 28, 2025 ▶ 21:50 From Idea to $650M Exit: Lessons in Building AI Startups · Y Combinator
Agarwal: Rigorous eval suites are core IP for leading AI startups
“The best AI companies will always have to be the edge of what the models can do, right? I think you always want to be threading the line. If everything works all the time, then you're not really pushing the limit and you're not innovating, right? So I think yo…”
Anish Agarwal Oct 5, 2025 ▶ 35:31 ⚡️Traversal: Causal ML and Reinforcement Learning
Husain: AI builders consistently get stuck moving demos to production
“Anytime that I try to help someone build an AI application, they always get stuck on how to move beyond a demo product. And they get stuck like how to systematically improve things and measure it.”
Hamel Husain Mar 13, 2025 ▶ 1:30 [Lightning Pod] Evals: How to Improve AI Consistently — with Hamel Husain and Shreya Shankar
Nguyen: AI bottleneck is evaluations rather than data as benchmarks saturate
“We are actually getting saturated in all benchmarks. So I think the bottleneck is actually in evaluations that we don't have all the frontier, like evals, like, I don't know GPGA, which is, like, A Google-proof question answering, like, PhD-level intelligence …”
Karina Nguyen Feb 9, 2025 ▶ 10:49 OpenAI researcher on why soft skills are the future of work | Karina Nguyen
Nguyen: AI model training requires evals where prompted baselines fail
“Prototype was prompted baseline. It's all, all, everything starts with, like, prompted baseline, and then, like, we craft, like, certain, like, evaluations that we want to, like, capture, that we want to, like, measure progress, at least, for the model, and th…”
Karina Nguyen Feb 1, 2025 ▶ 44:37 The Agent Reasoning Interface: Claude, ChatGPT Canvas, Tasks, Operator — with Karina Nguyen, OpenAI
MAD Insight
Evaluation is the single biggest bottleneck holding back enterprise AI adoption
“So, so I do think that evaluation is the biggest bottleneck for AI adoptions, because unless, like, if we can, like, if we can, like, develop a more reliable way to evaluate the application, that application is not going to get adopted. Like, or maybe, maybe i…”
Chip Huyen Jan 16, 2025 ▶ 35:43 What You MUST Know About AI Engineering | Chip Huyen, Author of “AI Engineering”

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.