evals

also referred to as: eval

17 statements across 8 episodes · 10 bullish · 3 bearish · 9 people on the record · first statement Aug 31, 2023 by Eugene Cheah · across every show →

Everything said about evals, oldest first

Aug 31, 2023 positive
Disclosure
RWKV Prioritizes User Feedback Over Benchmark Evals for Dataset Additions
“The reason why we add things to the data set was never about improving evals. It's about directly in response to user feedback.”
Eugene Cheah Aug 31, 2023 ▶ 43:11 RWKV: Reinventing RNNs for the Transformer Era
Aug 31, 2023 neutral
Insight
Foreign Language Data Degrades English Benchmark Scores on Small LLMs
“Adding in a foreign data set is actually a loss, because once you're below a certain param count, so we're talking about the seven important, right? The more you add that's more in line with your evals, the more it will degrade, and they just exclude it.”
Eugene Cheah Aug 31, 2023 ▶ 25:13 RWKV: Reinventing RNNs for the Transformer Era
Sep 17, 2024 positive
Insight
Building custom evals is high leverage for AI application developers
“I think for customers, and we work with a lot of customers, really developing their own evals is super high leverage. Because then you can upgrade really quickly when we have a new model, you can experiment with these things with confidence.”
Michelle Pokrass Sep 17, 2024 ▶ 37:26 Building AGI with OpenAI's Structured Outputs API
Oct 11, 2024 positive
Insight
Goyal: Write prompt evaluations before tweaking prompt text to measure impact
“The idea is like, it's useful to write the eval before you actually like tweak the prompt so that you can measure the impact of the tweak.”
Ankur Goyal Oct 11, 2024 ▶ 44:15 Production AI Engineering starts with Evals
Oct 11, 2024 positive
Insight
Goyal: Adopting evals resolves engineering stalemates over prompt and model choices
“And I think in the absence of evals, what I saw at Impera, and I see with almost all of our Customers before they start using brain trust is this kind of like stalemate between people on which prompt to use or which model to use or which technique to use that …”
Ankur Goyal Oct 11, 2024 ▶ 31:34 Production AI Engineering starts with Evals
Mar 13, 2025 positive
Insight
Shankar: Evals are necessary to train AI reasoning models
“You need evals to train your reasoning models.”
Shreya Shankar Mar 13, 2025 ▶ 27:13 [Lightning Pod] Evals: How to Improve AI Consistently — with Hamel Husain and Shreya Shankar
Mar 13, 2025
Insight
Husain: Teams rarely associate AI underperformance with a lack of evals
“Cause like one thing that I wrestle with is like evals is a solution, but the problem is, okay, your AI doesn't work, or it doesn't work as well as you want it to. And people don't associate the solution with the problem cleanly enough. Cause they don't know. …”
Hamel Husain Mar 13, 2025 ▶ 23:49 [Lightning Pod] Evals: How to Improve AI Consistently — with Hamel Husain and Shreya Shankar
Mar 13, 2025 negative
Opinion
Husain: 80% of LLM-as-a-judge implementations are unhelpful
“I feel like a 75% LMS judge because it's low effort is kind of easy, but I would say out of the 75%, 80% is not helpful.”
Hamel Husain Mar 13, 2025 ▶ 14:58 [Lightning Pod] Evals: How to Improve AI Consistently — with Hamel Husain and Shreya Shankar
May 23, 2025 positive
Prediction Not checkable as stated
Will Brown: Academia Will Likely Be the Best Source of AI Evals
“I mean, I do think that like the best source of evals going forward is probably going to be academia.”
Will Brown May 23, 2025 ▶ 23:43 ⚡️Multi-Turn RL for Multi-Hour Agents — with Will Brown, Prime Intellect
Oct 24, 2025 bearish
Opinion
Webster: AI evaluation tools are table-stakes commodities facing a feature-parity bloodbath
“I think evals are our table stakes. I think that they're a commodity and everyone should be doing them. And yes, there are companies that are doing great in the eval space, but To me, it just seemed like a bloodbath, you know, like we would just be, had a grea…”
Ian Webster Oct 24, 2025 ▶ 6:13 Breaking AI to Fix It: Ian Webster's Journey from Discord's Clyde to Promptfoo's $18M Series A
Dec 7, 2025 positive
Insight
Goyal: North Star AI Evals Prevent Test Brittleness
“Like if you construct evals in a way that represent the true north star of the problem that you're solving, then they tend not to break as you change the underlying system. Whereas if you hard code them to a narrow subset of like an implementation detail of yo…”
Ankur Goyal Dec 7, 2025 ▶ 13:47 The Great Evals Debate — Ankur Goyal & Malte Ubl
Dec 7, 2025 neutral
Insight
Goyal: Publishing Public Benchmarks Is Marketing, Not Product Improvement
“It's just that the value proposition of publishing an eval is completely orthogonal to the value proposition of building evals in service of building a good product. I think the purpose of publishing benchmarks is marketing, and it's good marketing.”
Ankur Goyal Dec 7, 2025 ▶ 7:53 The Great Evals Debate — Ankur Goyal & Malte Ubl
Dec 7, 2025 positive
Insight
Ubl: Evals Function to Tell Developers Overnight Whether a Change Is Good
“The way I think about evals is essentially like, it's the thing that, that can tell me tomorrow whether my change is good. And I can operate without that knowledge, but it's super, super helpful.”
Malte Ubl Dec 7, 2025 ▶ 2:28 The Great Evals Debate — Ankur Goyal & Malte Ubl
Dec 7, 2025 neutral
Assertion Not checkable as stated
Goyal: Commercial AI customers are reticent to give eval data to labs
“The interesting thing is that most customers, or actually I'd say a stronger statement, like all customers are quite afraid and reticent to just hand over the data that they use to do evals on to labs.”
Ankur Goyal Dec 7, 2025 ▶ 29:24 The Great Evals Debate — Ankur Goyal & Malte Ubl
Dec 7, 2025 positive
Insight
Malte Ubl: When Vibes and Eval Data Disagree, Vibes Are Right
“I think that the common quip that if the vibes and the data disagree, the vibes are probably right. It's true, right? So you have to like, be honest with yourself, like, do they agree and kind of iterate On them over time.”
Malte Ubl Dec 7, 2025 ▶ 12:13 The Great Evals Debate — Ankur Goyal & Malte Ubl
Dec 7, 2025 positive
Insight
Goyal: Providing eval criteria and examples is more effective than writing specs
“In many ways coming to the table of product building with representative examples and criteria that articulate what good versus bad is for a use case is just a more precise and usable form of product management than writing a spec.”
Ankur Goyal Dec 7, 2025 ▶ 18:36 The Great Evals Debate — Ankur Goyal & Malte Ubl
May 7, 2026 negative
Opinion
Pocock: Software developers are generally not interested in AI evaluations
“People are not really interested in evals, you know, like evals are not sexy. Like no one's excited to do evals these days, right?”
Matt Pocock May 7, 2026 ▶ 17:07 Senior Dev: This "Grill Me" Prompt Is Going Viral Among Top Engineers
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.