evals
also referred to as: eval
20 statements across 10 episodes · 8 bullish · 1 bearish · 13 people on the record · first statement Feb 9, 2025 by Karina Nguyen · across every show →
Everything said about evals, oldest first
Feb 9, 2025 neutral
Nguyen: Optimizing AI models constantly causes capability regressions across all labs
“If you optimize the model for this behavior, like, you kind of don't want to, like, brain damage in, like, other areas of intelligence, or, and this is happening, like, all the time in every lab and every, like, research team.”
Aug 9, 2025 positive
Turley: Evals are the lingua franca between product managers and AI researchers
“I was like, wow, this might be the lingua franca of how to communicate what the product should be doing. To people who do AI research. And that really clicked for me. And at the end of the day, it's not that different from The wisdom of you ought to articulate…”
Aug 31, 2025
Liu: Novel AI Product Discovery Should Start with Vibes, Not Evals
“For a completely novel product experience or form factor, you should actually not start with evals and you start with vibes, right? Meaning like, you know, you need to go and just kind of test in a much more open-ended way. Like, does this even work? Like, you…”
Sep 7, 2025 positive
Ezinne Udezue: AI PMs must master evaluations, not just prompt engineering
“There's this skill of being able to write evals. I know everybody can write prompts, prompt engineering. You can try and focus the LLM so that it can offer better insights and offer better results. But even as your LLM actually Provide, produces results. You n…”
Sep 18, 2025 neutral
Foody: Success measurement bottlenecks economy-wide AI automation
“And so in many ways, the barrier to applying agents to the entire economy To automate every workflow is how do we measure success? How do we eval it and write the PRDs for everything that we want agents to do, which Mercore is obviously a huge part of doing.”
Sep 18, 2025 positive
Sep 18, 2025 neutral
Sep 18, 2025 positive
Sep 25, 2025 neutral
Shreya Shankar: AI companies conceal evals because they are competitive moats
“And people don't talk about it because this is their moat, right? So people are not going to go and share all of these things because it makes sense, right? If you are an email writing assistant and you're doing this and you're doing it well, you don't want so…”
Sep 25, 2025 neutral
Husain: AI evals are just standard data science applied to AI products
“People say the word eval is trying to kind of like carve out this new thing, and saying, you know, evals, and then A-B testing, but if you zoom out, it's the same data science as before, and I think that's what's causing the confusion is, hey, we need data sci…”
Sep 25, 2025 neutral
Sep 25, 2025 positive
Rachitsky: Automated eval judges are the purest form of modern PRDs
“I've had some guests on the podcast recently who've been saying evals are the new PRDs. And if you look at this is exactly what this is like. Product managers, product teams, right? Here's what the product should be. Here's all the requirements. Here's like th…”
Sep 25, 2025 neutral
Husain: Jumping straight to evals without error analysis derails AI products
“You want to usually ground yourself in your actual errors. You don't want to skip this step. And so the reason I'm kind of spending so much time on this is like, this is where people get lost. They go straight into evals. Like, let me just write some tests. An…”
Oct 9, 2025 neutral
Jan 11, 2026
Badam: Relying entirely on fixed evals without team testing fails
“I don't think like if anybody's coming and seeing that, like my, I have this Concrete set of evals that I can, like, bet my life on, and then I don't need to think about anything else. Like, it's not going to work, and every new model that we're going to launc…”
Jan 11, 2026 negative
Reganti: Terms Like Evals and Agents Suffer From Semantic Diffusion
“I think Martin Fowler at some point had this term called semantic diffusion back in The 2000 which kind of means that someone comes up with a term, everybody starts butchering it with their own definitions, and then you kind of lose the actual definition of it…”
Jan 11, 2026
Jan 11, 2026 bullish
Badam: Organizational knowledge from trial-and-error evals is the decisive AI moat
“And that kind of knowledge that you've built across the organization or across like your own experience, lived experiences. I feel that the, that pain is what translates into the mode of the company, right? This could be like a product of evals or like somethi…”
Apr 23, 2026 positive
Jul 26, 2026 positive