Offline Evals

topic on 2 shows · 3 statements across 2 episodes

Latent Space the Official SaaStr Podcast

3 statements about Offline Evals, every show

Goyal: Offline Evals Are Hardest to Build but Most Efficient Feedback Loop
“Offline evals, which I think is what a lot of the debate was about are both the most challenging to build feedback loop and also the most efficient once built. And then I think AB tests are a little bit less challenging to build and a little bit less efficient…”
Ankur Goyal Dec 7, 2025 ▶ 2:57 The Great Evals Debate — Ankur Goyal & Malte Ubl
Goyal: Creating Golden Datasets for AI Evals Is Wasted Effort
“People don't really want to create golden data sets. It's, I think it's often a wasted effort to the point that you're making. I think the best teams view offline evals as a mechanism of reconciling what they see in production with real users who are using the…”
Ankur Goyal Dec 7, 2025 ▶ 10:17 The Great Evals Debate — Ankur Goyal & Malte Ubl
SAASTR Opinion
GitHub CPO: 95% Offline Evaluation Scores Still Yield Bad AI Products
“Even if you get an offline 95, by the way, usually that means the product, when it gets to market, it's a really bad product, at least in AI world, in my opinion.”
Mario Rodriguez Jan 10, 2025 ▶ 26:57 Adding AI to SaaS: Inside the AI Product Strategies of Figma, Cloudflare, GitHub and Ramp

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.