Kavak Spends Equal Time and Resources on Evals as on Agents
“We spend about the same amount of time, engineer time, tokens, and money on building the evals than building the agents.”
Cherny: AI evals saturate and must be discarded every few generations
“I think evals, they outlive the harness a little bit, but not quite that much. Like, an eval might live for maybe one, two, three model generations, but nowadays the, you know, we're on the exponential. The model is improving so quickly, very often we just sat…”
Penn: Evals Have Replaced Traditional PRDs in AI Product Management
“We actually have a saying on the team of evals are the new PRDs. Cause in order to deliver that user value it's not that exact artifact that people used to write in the last like one to two decades. It's a new way of working.”
Sanyal: AI developers should build evaluation suites before building applications
“In fact, you build the evals before you build the app. That kind of becomes, that's why people say that, hey, evals is the new weapon for product managers because they are the ones who are defining the app's behavior.”
Field: Design AI evaluations are inherently non-verifiable and require human judgment
“And the evals we're running on stuff we're doing, stuff others are doing, like a lot of it is inherently non-verifiable. It's like, gotta be human judged at the end of the day. And you can set up more automatic methods for that, but you still have that human i…”
Isenberg: Using customer data for AI evaluations creates effective sales assets
“It's also like low key, a really good sales asset, because imagine telling a property manager, You know, we tested this on, you know, 50 of your old maintenance requests. It routed 42 correctly, flagged six of them for human review, and made two mistakes. Here…”
Brown: Poker bot creation is a superior AI reasoning evaluation
“I think it's a nice eval because there is very little open source code for making poker bots. And there's a lot of published essays, there's a lot of published papers on it, but you really have to reason through everything.”
Nadella: Public AI benchmarks are gamed; companies need private evaluations
“Most importantly, you'll have private evals because we know all the evals out there are good, interesting, But they're not really that critical at this point because they're all can be maxed. And so the point is each company will have its own private eval.”
Bhatawdekar: Gen AI Systems Require Observability Feedback Loops for Evals
“So when you're building Gen AI systems, you really want that feedback loop of observability that helps you build better evals, that helps you ship better AI.”
Bhatawdekar: Rigorous evals are existential for AI apps built with 'vibe coding'
“When you're building these intelligent agentic applications using Vibe Coding evals almost become existential. You know, that's the only way you have a high degree of confidence that what you've built is going to work well.”
Pocock: Software developers are generally not interested in AI evaluations
“People are not really interested in evals, you know, like evals are not sexy. Like no one's excited to do evals these days, right?”
Anthropic PM lead: Teams only need 10 great evals, not hundreds
“You don't need to build hundreds of evals for them to be useful. Just building 10 great evals is important for helping the team quantify what the goal is and what their progress towards it is and what they're missing.”
Harrison Chase: AI evals and prompt optimization are closely tied, unlike memory
“I guess evals and prompt optimization are pretty closely tied, but like evals and memory are actually not at all tied.”
Diana Hu: Getting Good Prompts Requires Test-Driven Development via Evals
“The way you get a good prompt is all test-driven, just like evals, right? In a sense, the test cases are your evals.”
Badam: Relying solely on either evals or production monitoring is inadequate
“So I feel devals are important. Production monitoring is important, but this notion of only one of them is going to solve things for you. That is completely dismissible in my opinion.”
Reganti: Terms Like Evals and Agents Suffer From Semantic Diffusion
“I think Martin Fowler at some point had this term called semantic diffusion back in The 2000 which kind of means that someone comes up with a term, everybody starts butchering it with their own definitions, and then you kind of lose the actual definition of it…”
Badam: Relying entirely on fixed evals without team testing fails
“I don't think like if anybody's coming and seeing that, like my, I have this Concrete set of evals that I can, like, bet my life on, and then I don't need to think about anything else. Like, it's not going to work, and every new model that we're going to launc…”
Badam: Organizational knowledge from trial-and-error evals is the decisive AI moat
“And that kind of knowledge that you've built across the organization or across like your own experience, lived experiences. I feel that the, that pain is what translates into the mode of the company, right? This could be like a product of evals or like somethi…”
Ubl: Evals Function to Tell Developers Overnight Whether a Change Is Good
“The way I think about evals is essentially like, it's the thing that, that can tell me tomorrow whether my change is good. And I can operate without that knowledge, but it's super, super helpful.”
Goyal: Publishing Public Benchmarks Is Marketing, Not Product Improvement
“It's just that the value proposition of publishing an eval is completely orthogonal to the value proposition of building evals in service of building a good product. I think the purpose of publishing benchmarks is marketing, and it's good marketing.”
Malte Ubl: When Vibes and Eval Data Disagree, Vibes Are Right
“I think that the common quip that if the vibes and the data disagree, the vibes are probably right. It's true, right? So you have to like, be honest with yourself, like, do they agree and kind of iterate On them over time.”
Goyal: North Star AI Evals Prevent Test Brittleness
“Like if you construct evals in a way that represent the true north star of the problem that you're solving, then they tend not to break as you change the underlying system. Whereas if you hard code them to a narrow subset of like an implementation detail of yo…”
Goyal: Providing eval criteria and examples is more effective than writing specs
“In many ways coming to the table of product building with representative examples and criteria that articulate what good versus bad is for a use case is just a more precise and usable form of product management than writing a spec.”
Goyal: Commercial AI customers are reticent to give eval data to labs
“The interesting thing is that most customers, or actually I'd say a stronger statement, like all customers are quite afraid and reticent to just hand over the data that they use to do evals on to labs.”
Webster: AI evaluation tools are table-stakes commodities facing a feature-parity bloodbath
“I think evals are our table stakes. I think that they're a commodity and everyone should be doing them. And yes, there are companies that are doing great in the eval space, but To me, it just seemed like a bloodbath, you know, like we would just be, had a grea…”
Scale AI's enterprise and government work primarily consists of model evaluations
“A lot of it's evals and within enterprise customers and government customers, it's mostly evals because somebody has got to establish the benchmark for like what good looks like.”
Husain: Jumping straight to evals without error analysis derails AI products
“You want to usually ground yourself in your actual errors. You don't want to skip this step. And so the reason I'm kind of spending so much time on this is like, this is where people get lost. They go straight into evals. Like, let me just write some tests. An…”
Rachitsky: Automated eval judges are the purest form of modern PRDs
“I've had some guests on the podcast recently who've been saying evals are the new PRDs. And if you look at this is exactly what this is like. Product managers, product teams, right? Here's what the product should be. Here's all the requirements. Here's like th…”
Shreya Shankar: AI products typically need only four to seven LLM evals
“For me, like, between four and seven. It's not that many, because a lot of the failure modes, as Hamill said earlier, can be fixed by just fixing your prompt.”
Shreya Shankar: AI companies conceal evals because they are competitive moats
“And people don't talk about it because this is their moat, right? So people are not going to go and share all of these things because it makes sense, right? If you are an email writing assistant and you're doing this and you're doing it well, you don't want so…”
Husain: AI evals are just standard data science applied to AI products
“People say the word eval is trying to kind of like carve out this new thing, and saying, you know, evals, and then A-B testing, but if you zoom out, it's the same data science as before, and I think that's what's causing the confusion is, hey, we need data sci…”
Foody: AI evals are the product requirement documents for models
“If the model is the product, then the eval is the product requirement document.”
Foody: Success measurement bottlenecks economy-wide AI automation
“And so in many ways, the barrier to applying agents to the entire economy To automate every workflow is how do we measure success? How do we eval it and write the PRDs for everything that we want agents to do, which Mercore is obviously a huge part of doing.”
Foody: AI labs and apps will use evals as sales collateral
“I think labs will increasingly use labs as well as application layer companies will increasingly use evals to demonstrate the capabilities of their models and their products.”
Foody: AI evals and RL environments share the exact same data type
“There's not actually a nuance in the data type. It's more just a different semantic way of what describing what it's being used for. But ultimately it's just some stasis point for like, how do you measure what good looks like?”
If an AI model is the product, its eval is the PRD
“And if we think about the model as the product, then the eval is the PRD. And so many people have been sort of just like, you know, vibe spending on AI without actually writing the PRD of what do they want to implement and how do they measure that it's going t…”
Wu: Enterprise AI Evals Must Be Built Bottom-Up by Operators
“And evals also, oftentimes, need to come up bottom up. Right? Because all of these things are kind of in people's heads, in the actual operator's heads. Like, it's actually very hard to have a top-down mandate of, like, you got, like, this is how the evals sho…”
Ezinne Udezue: AI PMs must master evaluations, not just prompt engineering
“There's this skill of being able to write evals. I know everybody can write prompts, prompt engineering. You can try and focus the LLM so that it can offer better insights and offer better results. But even as your LLM actually Provide, produces results. You n…”
Liu: Novel AI Product Discovery Should Start with Vibes, Not Evals
“For a completely novel product experience or form factor, you should actually not start with evals and you start with vibes, right? Meaning like, you know, you need to go and just kind of test in a much more open-ended way. Like, does this even work? Like, you…”
Bonatsos: AI evals and professional domain modeling are a massive long-term market
“And I think that's a very big market opportunity that's opening up right now. It will be going on, you know, for a long time because you can bring in entire new professions and model them that you couldn't do, you know, as well today.”
Ng: Systematic error analysis with evals separates top AI agent teams
“The single biggest differentiator that I see in the market is, does the team know how to drive a systematic error analysis process with evals? So you're building the agents by analyzing at any moment in time, what's working, what's not working, what do you imp…”
Turley: Evals are the lingua franca between product managers and AI researchers
“I was like, wow, this might be the lingua franca of how to communicate what the product should be doing. To people who do AI research. And that really clicked for me. And at the end of the day, it's not that different from The wisdom of you ought to articulate…”
Kim: Designing good evaluations is the best way to motivate AI researchers
“If you want to nerdside someone into working on something, you just need to make a good eval, and then people are going to be so happy to try to hill climb that.”
Intellectually fascinating technical problems efficiently drive peer-to-peer customer acquisition
“I think that people without us having to force them to or ask them to, they are just interested in solving the problem around evals and they find it intellectually fascinating on their own. And I think that leads to a lot of organic knowledge sharing. A lot of…”
Laskin: Most Contributors on OpenAI's o1 Paper Worked on Evals
“When you look at the model card for, let's say, the O-one paper that came out, I think, last year. If you look at the distribution of what most people worked on, on that paper, it was evals.”
Jared Friedman: Evals, not prompts, are the core data asset for AI startups
“Even though we've been saying this for a year or more now, Gary, I think it's still the case that like evals are the true crown jewel Like data asset for all of these companies. Like one reason that power help was willing to open source the prompt is they told…”
Garry Tan: Vertical AI moats require sitting with domain workers to build evals
“You can't get the evals unless you are sitting literally side by side with people who are doing X, Y, or Z knowledge work. You know, you need to sit next to the tractor sales regional manager and understand, well, you know, this person cares about, you know, t…”
Will Brown: Academia Will Likely Be the Best Source of AI Evals
“I mean, I do think that like the best source of evals going forward is probably going to be academia.”
Garry Tan: Proprietary evaluations, not foundation models, are the true AI moat
“I mean, I think you know, ultimately the model itself is not the moat. Like I think that the evals themselves are the moat.”
Foody: Model Improvement via RL Is Gated Entirely by Evaluation Benchmarks
“Reinforcement learning is becoming so effective that once you create evals, the models can learn them and how to you know, improve capabilities. And so for everything that we want alums to be good at, we need evals for those things.”
Foody: Eval creation could become the most common knowledge job globally
“It
Would not surprise me if that becomes the most common knowledge work job in the world.”
Foody: Superintelligence cannot be recognized without comprehensive human evals
“You don't even know that you have super intelligence without having evals for everything. Cause it's like, you sort of need to understand what is the human baseline and like, what is good? It's like grounded in this like understanding of human behavior.”
Foody: Knowledge work will shift from repetitive tasks to fixed-cost eval building
“It does seem structurally more efficient for work to trend away from the variable cost of like doing it repeatedly towards this fixed cost of how do we build out the evals and the processes for models to do this themselves.”
Husain: 80% of LLM-as-a-judge implementations are unhelpful
“I feel like a 75% LMS judge because it's low effort is kind of easy, but I would say out of the 75%, 80% is not helpful.”
Husain: Teams rarely associate AI underperformance with a lack of evals
“Cause like one thing that I wrestle with is like evals is a solution, but the problem is, okay, your AI doesn't work, or it doesn't work as well as you want it to. And people don't associate the solution with the problem cleanly enough. Cause they don't know. …”
Shankar: Evals are necessary to train AI reasoning models
“You need evals to train your reasoning models.”
Chamath: AI labs are overfitting models on benchmarks, making scores unreliable
“And the problem, the dirty little secret of these model makers is that these guys are so trained on the evals that they're overfitting. And all this overfitting basically makes it pretty unreliable.”
Hiremath: AI model evaluation definitionally requires human-created datasets
“I think a lot of it will be human data going forward, and I think a great example of this is evals, right? Evals for models definitionally have to be outside of model capability, right? In order to see whether model is doing well at a particular task, You need…”
Nguyen: Optimizing AI models constantly causes capability regressions across all labs
“If you optimize the model for this behavior, like, you kind of don't want to, like, brain damage in, like, other areas of intelligence, or, and this is happening, like, all the time in every lab and every, like, research team.”
Goyal: Adopting evals resolves engineering stalemates over prompt and model choices
“And I think in the absence of evals, what I saw at Impera, and I see with almost all of our Customers before they start using brain trust is this kind of like stalemate between people on which prompt to use or which model to use or which technique to use that …”