evals

also referred to as: eval

20 statements across 10 episodes · 8 bullish · 1 bearish · 13 people on the record · first statement Feb 9, 2025 by Karina Nguyen · across every show →

Everything said about evals, oldest first

Feb 9, 2025 neutral
Assertion Not checkable as stated
Nguyen: Optimizing AI models constantly causes capability regressions across all labs
“If you optimize the model for this behavior, like, you kind of don't want to, like, brain damage in, like, other areas of intelligence, or, and this is happening, like, all the time in every lab and every, like, research team.”
Karina Nguyen Feb 9, 2025 ▶ 24:20 OpenAI researcher on why soft skills are the future of work | Karina Nguyen
Aug 9, 2025 positive
Insight
Turley: Evals are the lingua franca between product managers and AI researchers
“I was like, wow, this might be the lingua franca of how to communicate what the product should be doing. To people who do AI research. And that really clicked for me. And at the end of the day, it's not that different from The wisdom of you ought to articulate…”
Nick Turley Aug 9, 2025 ▶ 1:15:13 Inside ChatGPT: The fastest growing product in history | Nick Turley (OpenAI)
Aug 31, 2025
Insight
Liu: Novel AI Product Discovery Should Start with Vibes, Not Evals
“For a completely novel product experience or form factor, you should actually not start with evals and you start with vibes, right? Meaning like, you know, you need to go and just kind of test in a much more open-ended way. Like, does this even work? Like, you…”
Howie Liu Aug 31, 2025 ▶ 1:04:11 How we restructured Airtable's entire org for AI | Howie Liu (co-founder and CEO)
Sep 7, 2025 positive
Insight
Ezinne Udezue: AI PMs must master evaluations, not just prompt engineering
“There's this skill of being able to write evals. I know everybody can write prompts, prompt engineering. You can try and focus the LLM so that it can offer better insights and offer better results. But even as your LLM actually Provide, produces results. You n…”
Ezinne Udezue Sep 7, 2025 ▶ 20:01 How AI is reshaping the product role | Oji and Ezinne Udezue
Sep 18, 2025 neutral
Insight
Foody: Success measurement bottlenecks economy-wide AI automation
“And so in many ways, the barrier to applying agents to the entire economy To automate every workflow is how do we measure success? How do we eval it and write the PRDs for everything that we want agents to do, which Mercore is obviously a huge part of doing.”
Brendan Foody Sep 18, 2025 ▶ 7:07 Why experts writing AI evals is creating the fastest-growing companies in history | Brendan Foody
Sep 18, 2025 positive
Prediction Not checkable as stated
Foody: AI labs and apps will use evals as sales collateral
“I think labs will increasingly use labs as well as application layer companies will increasingly use evals to demonstrate the capabilities of their models and their products.”
Brendan Foody Sep 18, 2025 ▶ 9:16 Why experts writing AI evals is creating the fastest-growing companies in history | Brendan Foody
Sep 18, 2025 neutral
Insight
Foody: AI evals and RL environments share the exact same data type
“There's not actually a nuance in the data type. It's more just a different semantic way of what describing what it's being used for. But ultimately it's just some stasis point for like, how do you measure what good looks like?”
Brendan Foody Sep 18, 2025 ▶ 14:25 Why experts writing AI evals is creating the fastest-growing companies in history | Brendan Foody
Sep 18, 2025 positive
Insight
Foody: AI evals are the product requirement documents for models
“If the model is the product, then the eval is the product requirement document.”
Brendan Foody Sep 18, 2025 ▶ 6:40 Why experts writing AI evals is creating the fastest-growing companies in history | Brendan Foody
Sep 25, 2025 neutral
Opinion
Shreya Shankar: AI companies conceal evals because they are competitive moats
“And people don't talk about it because this is their moat, right? So people are not going to go and share all of these things because it makes sense, right? If you are an email writing assistant and you're doing this and you're doing it well, you don't want so…”
Shreya Shankar Sep 25, 2025 ▶ 1:08:23 Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar
Sep 25, 2025 neutral
Insight
Husain: AI evals are just standard data science applied to AI products
“People say the word eval is trying to kind of like carve out this new thing, and saying, you know, evals, and then A-B testing, but if you zoom out, it's the same data science as before, and I think that's what's causing the confusion is, hey, we need data sci…”
Hamel Husain Sep 25, 2025 ▶ 1:19:26 Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar
Sep 25, 2025 neutral
Insight
Shreya Shankar: AI products typically need only four to seven LLM evals
“For me, like, between four and seven. It's not that many, because a lot of the failure modes, as Hamill said earlier, can be fixed by just fixing your prompt.”
Shreya Shankar Sep 25, 2025 ▶ 1:05:19 Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar
Sep 25, 2025 positive
Opinion
Rachitsky: Automated eval judges are the purest form of modern PRDs
“I've had some guests on the podcast recently who've been saying evals are the new PRDs. And if you look at this is exactly what this is like. Product managers, product teams, right? Here's what the product should be. Here's all the requirements. Here's like th…”
Lenny Rachitsky Sep 25, 2025 ▶ 1:01:04 Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar
Sep 25, 2025 neutral
Insight
Husain: Jumping straight to evals without error analysis derails AI products
“You want to usually ground yourself in your actual errors. You don't want to skip this step. And so the reason I'm kind of spending so much time on this is like, this is where people get lost. They go straight into evals. Like, let me just write some tests. An…”
Hamel Husain Sep 25, 2025 ▶ 47:00 Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar
Oct 9, 2025 neutral
Disclosure
Scale AI's enterprise and government work primarily consists of model evaluations
“A lot of it's evals and within enterprise customers and government customers, it's mostly evals because somebody has got to establish the benchmark for like what good looks like.”
Jason Droege Oct 9, 2025 ▶ 31:40 Scale AI CEO on Meta’s $14B deal, scaling Uber Eats to $80B, & what frontier labs are building next
Jan 11, 2026
Insight
Badam: Relying entirely on fixed evals without team testing fails
“I don't think like if anybody's coming and seeing that, like my, I have this Concrete set of evals that I can, like, bet my life on, and then I don't need to think about anything else. Like, it's not going to work, and every new model that we're going to launc…”
Kiriti Badam Jan 11, 2026 ▶ 44:25 Why most AI products fail: Lessons from 50+ AI deployments at OpenAI, Google & Amazon
Jan 11, 2026 negative
Insight
Reganti: Terms Like Evals and Agents Suffer From Semantic Diffusion
“I think Martin Fowler at some point had this term called semantic diffusion back in The 2000 which kind of means that someone comes up with a term, everybody starts butchering it with their own definitions, and then you kind of lose the actual definition of it…”
Aishwarya Reganti (Ash) Jan 11, 2026 ▶ 39:31 Why most AI products fail: Lessons from 50+ AI deployments at OpenAI, Google & Amazon
Jan 11, 2026
Insight
Badam: Relying solely on either evals or production monitoring is inadequate
“So I feel devals are important. Production monitoring is important, but this notion of only one of them is going to solve things for you. That is completely dismissible in my opinion.”
Kiriti Badam Jan 11, 2026 ▶ 37:49 Why most AI products fail: Lessons from 50+ AI deployments at OpenAI, Google & Amazon
Jan 11, 2026 bullish
Insight
Badam: Organizational knowledge from trial-and-error evals is the decisive AI moat
“And that kind of knowledge that you've built across the organization or across like your own experience, lived experiences. I feel that the, that pain is what translates into the mode of the company, right? This could be like a product of evals or like somethi…”
Kiriti Badam Jan 11, 2026 ▶ 1:13:41 Why most AI products fail: Lessons from 50+ AI deployments at OpenAI, Google & Amazon
Apr 23, 2026 positive
Insight
Anthropic PM lead: Teams only need 10 great evals, not hundreds
“You don't need to build hundreds of evals for them to be useful. Just building 10 great evals is important for helping the team quantify what the goal is and what their progress towards it is and what they're missing.”
Cat Wu Apr 23, 2026 ▶ 55:11 How Anthropic’s product team moves faster than anyone else | Cat Wu (Head of Product, Claude Code)
Jul 26, 2026 positive
Insight
Penn: Evals Have Replaced Traditional PRDs in AI Product Management
“We actually have a saying on the team of evals are the new PRDs. Cause in order to deliver that user value it's not that exact artifact that people used to write in the last like one to two decades. It's a new way of working.”
Dianne Penn Jul 26, 2026 ▶ 41:36 Why AI is going vertical (again) | Dianne Penn (Anthropic)
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.