evals
also referred to as: eval
17 statements across 8 episodes · 10 bullish · 3 bearish · 9 people on the record · first statement Aug 31, 2023 by Eugene Cheah · across every show →
Everything said about evals, oldest first
Aug 31, 2023 positive
Aug 31, 2023 neutral
Foreign Language Data Degrades English Benchmark Scores on Small LLMs
“Adding in a foreign data set is actually a loss, because once you're below a certain param count, so we're talking about the seven important, right? The more you add that's more in line with your evals, the more it will degrade, and they just exclude it.”
Sep 17, 2024 positive
Building custom evals is high leverage for AI application developers
“I think for customers, and we work with a lot of customers, really developing their own evals is super high leverage. Because then you can upgrade really quickly when we have a new model, you can experiment with these things with confidence.”
Oct 11, 2024 positive
Oct 11, 2024 positive
Goyal: Adopting evals resolves engineering stalemates over prompt and model choices
“And I think in the absence of evals, what I saw at Impera, and I see with almost all of our Customers before they start using brain trust is this kind of like stalemate between people on which prompt to use or which model to use or which technique to use that …”
Mar 13, 2025 positive
Mar 13, 2025
Husain: Teams rarely associate AI underperformance with a lack of evals
“Cause like one thing that I wrestle with is like evals is a solution, but the problem is, okay, your AI doesn't work, or it doesn't work as well as you want it to. And people don't associate the solution with the problem cleanly enough. Cause they don't know. …”
Mar 13, 2025 negative
May 23, 2025 positive
Oct 24, 2025 bearish
Webster: AI evaluation tools are table-stakes commodities facing a feature-parity bloodbath
“I think evals are our table stakes. I think that they're a commodity and everyone should be doing them. And yes, there are companies that are doing great in the eval space, but To me, it just seemed like a bloodbath, you know, like we would just be, had a grea…”
Dec 7, 2025 positive
Goyal: North Star AI Evals Prevent Test Brittleness
“Like if you construct evals in a way that represent the true north star of the problem that you're solving, then they tend not to break as you change the underlying system. Whereas if you hard code them to a narrow subset of like an implementation detail of yo…”
Dec 7, 2025 neutral
Goyal: Publishing Public Benchmarks Is Marketing, Not Product Improvement
“It's just that the value proposition of publishing an eval is completely orthogonal to the value proposition of building evals in service of building a good product. I think the purpose of publishing benchmarks is marketing, and it's good marketing.”
Dec 7, 2025 positive
Dec 7, 2025 neutral
Dec 7, 2025 positive
Dec 7, 2025 positive
Goyal: Providing eval criteria and examples is more effective than writing specs
“In many ways coming to the table of product building with representative examples and criteria that articulate what good versus bad is for a use case is just a more precise and usable form of product management than writing a spec.”