Evaluation
topic on 6 shows · 8 statements across 8 episodes
the Y Combinator Startup Podcast
Latent Space
Lenny's Podcast
the MAD Podcast
the a16z Podcast
20VC
8 statements about Evaluation, every show
Complex AI evals require smaller sample sizes and broader criteria
“Evaluations as they become more complex, Have a fewer sample size, but a larger set of criteria or expectations of them.”
Joelle Pineau: AI benchmarks should be treated as narrow unit tests
“Think of evaluations as, like, unit test for the performance of your system. I mean, software engineers will know what that is, right? Like, you run through that evaluation, and that gives you, like, a signal of how the system is doing in a particular dimensio…”
Heller: Casetext derived most evals from customer failures, not lab tests
“We've added much more evals at this point from real things that happened to real customers than the ones we came up with in the lab.”
Agarwal: Rigorous eval suites are core IP for leading AI startups
“The best AI companies will always have to be the edge of what the models can do, right? I think you always want to be threading the line. If everything works all the time, then you're not really pushing the limit and you're not innovating, right? So I think yo…”
Husain: AI builders consistently get stuck moving demos to production
“Anytime that I try to help someone build an AI application, they always get stuck on how to move beyond a demo product. And they get stuck like how to systematically improve things and measure it.”
Nguyen: AI bottleneck is evaluations rather than data as benchmarks saturate
“We are actually getting saturated in all benchmarks. So I think the bottleneck is actually in evaluations that we don't have all the frontier, like evals, like, I don't know GPGA, which is, like, A Google-proof question answering, like, PhD-level intelligence …”
Nguyen: AI model training requires evals where prompted baselines fail
“Prototype was prompted baseline. It's all, all, everything starts with, like, prompted baseline, and then, like, we craft, like, certain, like, evaluations that we want to, like, capture, that we want to, like, measure progress, at least, for the model, and th…”
Evaluation is the single biggest bottleneck holding back enterprise AI adoption
“So, so I do think that evaluation is the biggest bottleneck for AI adoptions, because unless, like, if we can, like, if we can, like, develop a more reliable way to evaluate the application, that application is not going to get adopted. Like, or maybe, maybe i…”