Eval Datasets

topic on 1 show · 2 statements across 1 episodes

the Neon Show

2 statements about Eval Datasets, every show

NEON SHOW Insight
Bhatawdekar: Production eval datasets must be continuously updated from live logs
“These eval datasets, they are not static. As the teams look at their logs, at how their systems are working in the real world, in the production use cases, they're able to leverage those insights to continually augment their eval datasets. And so these eval da…”
Ameya Bhatawdekar May 15, 2026 ▶ 8:51 The “Messy State” of AI & How to Fix It | Ameya, Braintrust CTO
NEON SHOW Assertion Not checkable as stated
Bhatawdekar: Notion, Stripe, and Zapier run automated evals on every system modification
“All of these companies are building agents, intelligent systems that are performing specialized tasks for whatever use cases they have, and so all of them build evals that reflect how their systems are expected to behave in production. They're all following si…”
Ameya Bhatawdekar May 15, 2026 ▶ 8:07 The “Messy State” of AI & How to Fix It | Ameya, Braintrust CTO

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.