evaluation

also referred to as: evaluations

3 statements across 3 episodes · 1 bullish · 0 bearish · 3 people on the record · first statement Feb 1, 2025 by Karina Nguyen · across every show →

Everything said about evaluation, oldest first

Feb 1, 2025
Insight
Nguyen: AI model training requires evals where prompted baselines fail
“Prototype was prompted baseline. It's all, all, everything starts with, like, prompted baseline, and then, like, we craft, like, certain, like, evaluations that we want to, like, capture, that we want to, like, measure progress, at least, for the model, and th…”
Karina Nguyen Feb 1, 2025 ▶ 44:37 The Agent Reasoning Interface: Claude, ChatGPT Canvas, Tasks, Operator — with Karina Nguyen, OpenAI
Mar 13, 2025
Insight
Husain: AI builders consistently get stuck moving demos to production
“Anytime that I try to help someone build an AI application, they always get stuck on how to move beyond a demo product. And they get stuck like how to systematically improve things and measure it.”
Hamel Husain Mar 13, 2025 ▶ 1:30 [Lightning Pod] Evals: How to Improve AI Consistently — with Hamel Husain and Shreya Shankar
Oct 5, 2025 positive
Insight
Agarwal: Rigorous eval suites are core IP for leading AI startups
“The best AI companies will always have to be the edge of what the models can do, right? I think you always want to be threading the line. If everything works all the time, then you're not really pushing the limit and you're not innovating, right? So I think yo…”
Anish Agarwal Oct 5, 2025 ▶ 35:31 ⚡️Traversal: Causal ML and Reinforcement Learning
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.