eval datasets

2 statements across 1 episodes · 2 bullish · 0 bearish · 1 people on the record · first statement May 15, 2026 by Ameya Bhatawdekar · across every show →

Everything said about eval datasets, oldest first

May 15, 2026 positive
Insight
Bhatawdekar: Production eval datasets must be continuously updated from live logs
“These eval datasets, they are not static. As the teams look at their logs, at how their systems are working in the real world, in the production use cases, they're able to leverage those insights to continually augment their eval datasets. And so these eval da…”
Ameya Bhatawdekar May 15, 2026 ▶ 8:51 The “Messy State” of AI & How to Fix It | Ameya, Braintrust CTO
May 15, 2026 positive
Assertion Not checkable as stated
Bhatawdekar: Notion, Stripe, and Zapier run automated evals on every system modification
“All of these companies are building agents, intelligent systems that are performing specialized tasks for whatever use cases they have, and so all of them build evals that reflect how their systems are expected to behave in production. They're all following si…”
Ameya Bhatawdekar May 15, 2026 ▶ 8:07 The “Messy State” of AI & How to Fix It | Ameya, Braintrust CTO
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.