Evals

topic on 15 shows · 66 statements across 42 episodes

the Y Combinator Startup Podcast In Depth BG2 Pod the Knowledge Project Latent Space Lenny's Podcast the Neon Show No Priors the Startup Ideas Podcast WTF is with Nikhil Kamath Sourcery the MAD Podcast the a16z Podcast All-In 20VC

The latest 60 statements about Evals, every show

a16z Disclosure
Kavak Spends Equal Time and Resources on Evals as on Agents
“We spend about the same amount of time, engineer time, tokens, and money on building the evals than building the agents.”
Ale Massa Aug 9, 2026 ▶ 9:29 Kavak's Playbook for Rebuilding a Company Around AI
Cherny: AI evals saturate and must be discarded every few generations
“I think evals, they outlive the harness a little bit, but not quite that much. Like, an eval might live for maybe one, two, three model generations, but nowadays the, you know, we're on the exponential. The model is improving so quickly, very often we just sat…”
Boris Cherny Jul 27, 2026 ▶ 10:01 Boris Cherny: We Cut 80% of Claude Code’s Prompt · Y Combinator
Penn: Evals Have Replaced Traditional PRDs in AI Product Management
“We actually have a saying on the team of evals are the new PRDs. Cause in order to deliver that user value it's not that exact artifact that people used to write in the last like one to two decades. It's a new way of working.”
Dianne Penn Jul 26, 2026 ▶ 41:36 Why AI is going vertical (again) | Dianne Penn (Anthropic)
NEON SHOW Insight
Sanyal: AI developers should build evaluation suites before building applications
“In fact, you build the evals before you build the app. That kind of becomes, that's why people say that, hey, evals is the new weapon for product managers because they are the ones who are defining the app's behavior.”
Atin Sanyal Jul 9, 2026 ▶ 45:24 The Hidden Layer Every Al Product Needs Today | Atin Sanyal Founder, Galileo
SOURCERY Insight
Field: Design AI evaluations are inherently non-verifiable and require human judgment
“And the evals we're running on stuff we're doing, stuff others are doing, like a lot of it is inherently non-verifiable. It's like, gotta be human judged at the end of the day. And you can set up more automatic methods for that, but you still have that human i…”
Dylan Field Jul 2, 2026 ▶ 9:59 Dylan Field on the “Permanent Underclass of Zero Taste” · Sourcery with Molly O'Shea
Isenberg: Using customer data for AI evaluations creates effective sales assets
“It's also like low key, a really good sales asset, because imagine telling a property manager, You know, we tested this on, you know, 50 of your old maintenance requests. It routed 42 correctly, flagged six of them for human review, and made two mistakes. Here…”
Greg Isenberg Jul 1, 2026 ▶ 15:12 AI Agents are the new SaaS
NO PRIORS Insight
Brown: Poker bot creation is a superior AI reasoning evaluation
“I think it's a nice eval because there is very little open source code for making poker bots. And there's a lot of published essays, there's a lot of published papers on it, but you really have to reason through everything.”
Noam Brown Jun 26, 2026 ▶ 8:41 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
NO PRIORS Insight
Nadella: Public AI benchmarks are gamed; companies need private evaluations
“Most importantly, you'll have private evals because we know all the evals out there are good, interesting, But they're not really that critical at this point because they're all can be maxed. And so the point is each company will have its own private eval.”
Satya Nadella Jun 4, 2026 ▶ 4:54 We Need An Ecosystem in AI, And Every Company Can Win A Place In It
NEON SHOW Insight
Bhatawdekar: Gen AI Systems Require Observability Feedback Loops for Evals
“So when you're building Gen AI systems, you really want that feedback loop of observability that helps you build better evals, that helps you ship better AI.”
Ameya Bhatawdekar May 15, 2026 ▶ 4:49 The “Messy State” of AI & How to Fix It | Ameya, Braintrust CTO
NEON SHOW Insight
Bhatawdekar: Rigorous evals are existential for AI apps built with 'vibe coding'
“When you're building these intelligent agentic applications using Vibe Coding evals almost become existential. You know, that's the only way you have a high degree of confidence that what you've built is going to work well.”
Ameya Bhatawdekar May 15, 2026 ▶ 24:53 The “Messy State” of AI & How to Fix It | Ameya, Braintrust CTO
Pocock: Software developers are generally not interested in AI evaluations
“People are not really interested in evals, you know, like evals are not sexy. Like no one's excited to do evals these days, right?”
Matt Pocock May 7, 2026 ▶ 17:07 Senior Dev: This "Grill Me" Prompt Is Going Viral Among Top Engineers
Anthropic PM lead: Teams only need 10 great evals, not hundreds
“You don't need to build hundreds of evals for them to be useful. Just building 10 great evals is important for helping the team quantify what the goal is and what their progress towards it is and what they're missing.”
Cat Wu Apr 23, 2026 ▶ 55:11 How Anthropic’s product team moves faster than anyone else | Cat Wu (Head of Product, Claude Code)
MAD Insight
Harrison Chase: AI evals and prompt optimization are closely tied, unlike memory
“I guess evals and prompt optimization are pretty closely tied, but like evals and memory are actually not at all tied.”
Harrison Chase Mar 12, 2026 ▶ 43:25 Everything Gets Rebuilt: The New AI Agent Stack | Harrison Chase, LangChain
Diana Hu: Getting Good Prompts Requires Test-Driven Development via Evals
“The way you get a good prompt is all test-driven, just like evals, right? In a sense, the test cases are your evals.”
Diana Hu Feb 6, 2026 ▶ 36:30 We're All Addicted To Claude Code · Y Combinator
Badam: Relying solely on either evals or production monitoring is inadequate
“So I feel devals are important. Production monitoring is important, but this notion of only one of them is going to solve things for you. That is completely dismissible in my opinion.”
Kiriti Badam Jan 11, 2026 ▶ 37:49 Why most AI products fail: Lessons from 50+ AI deployments at OpenAI, Google & Amazon
Reganti: Terms Like Evals and Agents Suffer From Semantic Diffusion
“I think Martin Fowler at some point had this term called semantic diffusion back in The 2000 which kind of means that someone comes up with a term, everybody starts butchering it with their own definitions, and then you kind of lose the actual definition of it…”
Aishwarya Reganti (Ash) Jan 11, 2026 ▶ 39:31 Why most AI products fail: Lessons from 50+ AI deployments at OpenAI, Google & Amazon
Badam: Relying entirely on fixed evals without team testing fails
“I don't think like if anybody's coming and seeing that, like my, I have this Concrete set of evals that I can, like, bet my life on, and then I don't need to think about anything else. Like, it's not going to work, and every new model that we're going to launc…”
Kiriti Badam Jan 11, 2026 ▶ 44:25 Why most AI products fail: Lessons from 50+ AI deployments at OpenAI, Google & Amazon
Badam: Organizational knowledge from trial-and-error evals is the decisive AI moat
“And that kind of knowledge that you've built across the organization or across like your own experience, lived experiences. I feel that the, that pain is what translates into the mode of the company, right? This could be like a product of evals or like somethi…”
Kiriti Badam Jan 11, 2026 ▶ 1:13:41 Why most AI products fail: Lessons from 50+ AI deployments at OpenAI, Google & Amazon
Ubl: Evals Function to Tell Developers Overnight Whether a Change Is Good
“The way I think about evals is essentially like, it's the thing that, that can tell me tomorrow whether my change is good. And I can operate without that knowledge, but it's super, super helpful.”
Malte Ubl Dec 7, 2025 ▶ 2:28 The Great Evals Debate — Ankur Goyal & Malte Ubl
Goyal: Publishing Public Benchmarks Is Marketing, Not Product Improvement
“It's just that the value proposition of publishing an eval is completely orthogonal to the value proposition of building evals in service of building a good product. I think the purpose of publishing benchmarks is marketing, and it's good marketing.”
Ankur Goyal Dec 7, 2025 ▶ 7:53 The Great Evals Debate — Ankur Goyal & Malte Ubl
Malte Ubl: When Vibes and Eval Data Disagree, Vibes Are Right
“I think that the common quip that if the vibes and the data disagree, the vibes are probably right. It's true, right? So you have to like, be honest with yourself, like, do they agree and kind of iterate On them over time.”
Malte Ubl Dec 7, 2025 ▶ 12:13 The Great Evals Debate — Ankur Goyal & Malte Ubl
Goyal: North Star AI Evals Prevent Test Brittleness
“Like if you construct evals in a way that represent the true north star of the problem that you're solving, then they tend not to break as you change the underlying system. Whereas if you hard code them to a narrow subset of like an implementation detail of yo…”
Ankur Goyal Dec 7, 2025 ▶ 13:47 The Great Evals Debate — Ankur Goyal & Malte Ubl
Goyal: Providing eval criteria and examples is more effective than writing specs
“In many ways coming to the table of product building with representative examples and criteria that articulate what good versus bad is for a use case is just a more precise and usable form of product management than writing a spec.”
Ankur Goyal Dec 7, 2025 ▶ 18:36 The Great Evals Debate — Ankur Goyal & Malte Ubl
LATENT SPACE Assertion Not checkable as stated
Goyal: Commercial AI customers are reticent to give eval data to labs
“The interesting thing is that most customers, or actually I'd say a stronger statement, like all customers are quite afraid and reticent to just hand over the data that they use to do evals on to labs.”
Ankur Goyal Dec 7, 2025 ▶ 29:24 The Great Evals Debate — Ankur Goyal & Malte Ubl
Webster: AI evaluation tools are table-stakes commodities facing a feature-parity bloodbath
“I think evals are our table stakes. I think that they're a commodity and everyone should be doing them. And yes, there are companies that are doing great in the eval space, but To me, it just seemed like a bloodbath, you know, like we would just be, had a grea…”
Ian Webster Oct 24, 2025 ▶ 6:13 Breaking AI to Fix It: Ian Webster's Journey from Discord's Clyde to Promptfoo's $18M Series A
Scale AI's enterprise and government work primarily consists of model evaluations
“A lot of it's evals and within enterprise customers and government customers, it's mostly evals because somebody has got to establish the benchmark for like what good looks like.”
Jason Droege Oct 9, 2025 ▶ 31:40 Scale AI CEO on Meta’s $14B deal, scaling Uber Eats to $80B, & what frontier labs are building next
Husain: Jumping straight to evals without error analysis derails AI products
“You want to usually ground yourself in your actual errors. You don't want to skip this step. And so the reason I'm kind of spending so much time on this is like, this is where people get lost. They go straight into evals. Like, let me just write some tests. An…”
Hamel Husain Sep 25, 2025 ▶ 47:00 Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar
Rachitsky: Automated eval judges are the purest form of modern PRDs
“I've had some guests on the podcast recently who've been saying evals are the new PRDs. And if you look at this is exactly what this is like. Product managers, product teams, right? Here's what the product should be. Here's all the requirements. Here's like th…”
Lenny Rachitsky Sep 25, 2025 ▶ 1:01:04 Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar
Shreya Shankar: AI products typically need only four to seven LLM evals
“For me, like, between four and seven. It's not that many, because a lot of the failure modes, as Hamill said earlier, can be fixed by just fixing your prompt.”
Shreya Shankar Sep 25, 2025 ▶ 1:05:19 Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar
Shreya Shankar: AI companies conceal evals because they are competitive moats
“And people don't talk about it because this is their moat, right? So people are not going to go and share all of these things because it makes sense, right? If you are an email writing assistant and you're doing this and you're doing it well, you don't want so…”
Shreya Shankar Sep 25, 2025 ▶ 1:08:23 Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar
Husain: AI evals are just standard data science applied to AI products
“People say the word eval is trying to kind of like carve out this new thing, and saying, you know, evals, and then A-B testing, but if you zoom out, it's the same data science as before, and I think that's what's causing the confusion is, hey, we need data sci…”
Hamel Husain Sep 25, 2025 ▶ 1:19:26 Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar
Foody: AI evals are the product requirement documents for models
“If the model is the product, then the eval is the product requirement document.”
Brendan Foody Sep 18, 2025 ▶ 6:40 Why experts writing AI evals is creating the fastest-growing companies in history | Brendan Foody
Foody: Success measurement bottlenecks economy-wide AI automation
“And so in many ways, the barrier to applying agents to the entire economy To automate every workflow is how do we measure success? How do we eval it and write the PRDs for everything that we want agents to do, which Mercore is obviously a huge part of doing.”
Brendan Foody Sep 18, 2025 ▶ 7:07 Why experts writing AI evals is creating the fastest-growing companies in history | Brendan Foody
LENNY'S PODCAST Prediction Not checkable as stated
Foody: AI labs and apps will use evals as sales collateral
“I think labs will increasingly use labs as well as application layer companies will increasingly use evals to demonstrate the capabilities of their models and their products.”
Brendan Foody Sep 18, 2025 ▶ 9:16 Why experts writing AI evals is creating the fastest-growing companies in history | Brendan Foody
Foody: AI evals and RL environments share the exact same data type
“There's not actually a nuance in the data type. It's more just a different semantic way of what describing what it's being used for. But ultimately it's just some stasis point for like, how do you measure what good looks like?”
Brendan Foody Sep 18, 2025 ▶ 14:25 Why experts writing AI evals is creating the fastest-growing companies in history | Brendan Foody
20VC Insight
If an AI model is the product, its eval is the PRD
“And if we think about the model as the product, then the eval is the PRD. And so many people have been sort of just like, you know, vibe spending on AI without actually writing the PRD of what do they want to implement and how do they measure that it's going t…”
Brendan Foody Sep 15, 2025 ▶ 35:15 Mercor CEO & Co-Founder, Brendan Foody: How They Grew from $1M to $500M in 17 Months · 20VC with Harry Stebbings
BG2 Insight
Wu: Enterprise AI Evals Must Be Built Bottom-Up by Operators
“And evals also, oftentimes, need to come up bottom up. Right? Because all of these things are kind of in people's heads, in the actual operator's heads. Like, it's actually very hard to have a top-down mandate of, like, you got, like, this is how the evals sho…”
Sherwin Wu Sep 11, 2025 ▶ 19:17 Inside OpenAI Enterprise: Forward Deployed Engineering, GPT-5, and More | BG2 Guest Interview · Bg2 Pod
Ezinne Udezue: AI PMs must master evaluations, not just prompt engineering
“There's this skill of being able to write evals. I know everybody can write prompts, prompt engineering. You can try and focus the LLM so that it can offer better insights and offer better results. But even as your LLM actually Provide, produces results. You n…”
Ezinne Udezue Sep 7, 2025 ▶ 20:01 How AI is reshaping the product role | Oji and Ezinne Udezue
Liu: Novel AI Product Discovery Should Start with Vibes, Not Evals
“For a completely novel product experience or form factor, you should actually not start with evals and you start with vibes, right? Meaning like, you know, you need to go and just kind of test in a much more open-ended way. Like, does this even work? Like, you…”
Howie Liu Aug 31, 2025 ▶ 1:04:11 How we restructured Airtable's entire org for AI | Howie Liu (co-founder and CEO)
WTF Prediction Not checkable as stated
Bonatsos: AI evals and professional domain modeling are a massive long-term market
“And I think that's a very big market opportunity that's opening up right now. It will be going on, you know, for a long time because you can bring in entire new professions and model them that you couldn't do, you know, as well today.”
Niko Bonatsos Aug 29, 2025 ▶ 31:07 Inside Silicon Valley’s VC Playbook | WTF is Venture Capital? - 2025 Edition | Ep. 24 · Nikhil Kamath
NO PRIORS Insight
Ng: Systematic error analysis with evals separates top AI agent teams
“The single biggest differentiator that I see in the market is, does the team know how to drive a systematic error analysis process with evals? So you're building the agents by analyzing at any moment in time, what's working, what's not working, what do you imp…”
Andrew Ng Aug 21, 2025 ▶ 3:29 No Priors Ep. 128 | With Andrew Ng, Managing General Partner at AI Fund
Turley: Evals are the lingua franca between product managers and AI researchers
“I was like, wow, this might be the lingua franca of how to communicate what the product should be doing. To people who do AI research. And that really clicked for me. And at the end of the day, it's not that different from The wisdom of you ought to articulate…”
Nick Turley Aug 9, 2025 ▶ 1:15:13 Inside ChatGPT: The fastest growing product in history | Nick Turley (OpenAI)
a16z Insight
Kim: Designing good evaluations is the best way to motivate AI researchers
“If you want to nerdside someone into working on something, you just need to make a good eval, and then people are going to be so happy to try to hill climb that.”
Christina Kim Aug 8, 2025 ▶ 11:07 GPT-5 and Agents Breakdown – w/ OpenAI Researchers Isa Fulford & Christina Kim
IN DEPTH Insight
Intellectually fascinating technical problems efficiently drive peer-to-peer customer acquisition
“I think that people without us having to force them to or ask them to, they are just interested in solving the problem around evals and they find it intellectually fascinating on their own. And I think that leads to a lot of organic knowledge sharing. A lot of…”
Ankur Goyal Jul 24, 2025 ▶ 33:31 What Braintrust got right about product-market fit | Ankur Goyal (Founder and CEO)
NO PRIORS Assertion Contradicted
Laskin: Most Contributors on OpenAI's o1 Paper Worked on Evals
“When you look at the model card for, let's say, the O-one paper that came out, I think, last year. If you look at the distribution of what most people worked on, on that paper, it was evals.”
Misha Laskin Jul 17, 2025 ▶ 17:16 No Priors Ep. 123 | With ReflectionAI Co-Founder and CEO Misha Laskin
Jared Friedman: Evals, not prompts, are the core data asset for AI startups
“Even though we've been saying this for a year or more now, Gary, I think it's still the case that like evals are the true crown jewel Like data asset for all of these companies. Like one reason that power help was willing to open source the prompt is they told…”
Jared Friedman May 30, 2025 ▶ 14:27 State-Of-The-Art Prompting For AI Agents · Y Combinator
Garry Tan: Vertical AI moats require sitting with domain workers to build evals
“You can't get the evals unless you are sitting literally side by side with people who are doing X, Y, or Z knowledge work. You know, you need to sit next to the tractor sales regional manager and understand, well, you know, this person cares about, you know, t…”
Garry Tan May 30, 2025 ▶ 15:02 State-Of-The-Art Prompting For AI Agents · Y Combinator
LATENT SPACE Prediction Not checkable as stated
Will Brown: Academia Will Likely Be the Best Source of AI Evals
“I mean, I do think that like the best source of evals going forward is probably going to be academia.”
Will Brown May 23, 2025 ▶ 23:43 ⚡️Multi-Turn RL for Multi-Hour Agents — with Will Brown, Prime Intellect
Garry Tan: Proprietary evaluations, not foundation models, are the true AI moat
“I mean, I think you know, ultimately the model itself is not the moat. Like I think that the evals themselves are the moat.”
Garry Tan Apr 29, 2025 ▶ 1:52:52 Y Combinator CEO Garry Tan: Turning Ambitious Misfits into Founders
NO PRIORS Insight
Foody: Model Improvement via RL Is Gated Entirely by Evaluation Benchmarks
“Reinforcement learning is becoming so effective that once you create evals, the models can learn them and how to you know, improve capabilities. And so for everything that we want alums to be good at, we need evals for those things.”
Brendan Foody Apr 10, 2025 ▶ 1:14 No Priors Ep. 110 | With Mercor CEO and Co-Founder Brendan Foody
NO PRIORS Prediction Not checkable as stated
Foody: Eval creation could become the most common knowledge job globally
“It Would not surprise me if that becomes the most common knowledge work job in the world.”
Brendan Foody Apr 10, 2025 ▶ 19:43 No Priors Ep. 110 | With Mercor CEO and Co-Founder Brendan Foody
NO PRIORS Insight
Foody: Superintelligence cannot be recognized without comprehensive human evals
“You don't even know that you have super intelligence without having evals for everything. Cause it's like, you sort of need to understand what is the human baseline and like, what is good? It's like grounded in this like understanding of human behavior.”
Brendan Foody Apr 10, 2025 ▶ 20:09 No Priors Ep. 110 | With Mercor CEO and Co-Founder Brendan Foody
NO PRIORS Insight
Foody: Knowledge work will shift from repetitive tasks to fixed-cost eval building
“It does seem structurally more efficient for work to trend away from the variable cost of like doing it repeatedly towards this fixed cost of how do we build out the evals and the processes for models to do this themselves.”
Brendan Foody Apr 10, 2025 ▶ 21:13 No Priors Ep. 110 | With Mercor CEO and Co-Founder Brendan Foody
Husain: 80% of LLM-as-a-judge implementations are unhelpful
“I feel like a 75% LMS judge because it's low effort is kind of easy, but I would say out of the 75%, 80% is not helpful.”
Hamel Husain Mar 13, 2025 ▶ 14:58 [Lightning Pod] Evals: How to Improve AI Consistently — with Hamel Husain and Shreya Shankar
Husain: Teams rarely associate AI underperformance with a lack of evals
“Cause like one thing that I wrestle with is like evals is a solution, but the problem is, okay, your AI doesn't work, or it doesn't work as well as you want it to. And people don't associate the solution with the problem cleanly enough. Cause they don't know. …”
Hamel Husain Mar 13, 2025 ▶ 23:49 [Lightning Pod] Evals: How to Improve AI Consistently — with Hamel Husain and Shreya Shankar
Shankar: Evals are necessary to train AI reasoning models
“You need evals to train your reasoning models.”
Shreya Shankar Mar 13, 2025 ▶ 27:13 [Lightning Pod] Evals: How to Improve AI Consistently — with Hamel Husain and Shreya Shankar
ALL-IN Assertion Not checkable as stated
Chamath: AI labs are overfitting models on benchmarks, making scores unreliable
“And the problem, the dirty little secret of these model makers is that these guys are so trained on the evals that they're overfitting. And all this overfitting basically makes it pretty unreliable.”
Chamath Palihapitiya Mar 8, 2025 ▶ 1:27:47 Tariffs, Trump's Economic Endgame, Market Chaos, Bitcoin Reserve, CoreWeave IPO
20VC Insight
Hiremath: AI model evaluation definitionally requires human-created datasets
“I think a lot of it will be human data going forward, and I think a great example of this is evals, right? Evals for models definitionally have to be outside of model capability, right? In order to see whether model is doing well at a particular task, You need…”
Adarsh Hiremath Feb 20, 2025 ▶ 18:49 Adarsh Hiremath @ Mercor: The Fastest Growing Startup in Silicon Valley | E1261 · 20VC with Harry Stebbings
LENNY'S PODCAST Assertion Not checkable as stated
Nguyen: Optimizing AI models constantly causes capability regressions across all labs
“If you optimize the model for this behavior, like, you kind of don't want to, like, brain damage in, like, other areas of intelligence, or, and this is happening, like, all the time in every lab and every, like, research team.”
Karina Nguyen Feb 9, 2025 ▶ 24:20 OpenAI researcher on why soft skills are the future of work | Karina Nguyen
Goyal: Adopting evals resolves engineering stalemates over prompt and model choices
“And I think in the absence of evals, what I saw at Impera, and I see with almost all of our Customers before they start using brain trust is this kind of like stalemate between people on which prompt to use or which model to use or which technique to use that …”
Ankur Goyal Oct 11, 2024 ▶ 31:34 Production AI Engineering starts with Evals

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.