benchmarks
7 statements across 6 episodes · 1 bullish · 3 bearish · 5 people on the record · first statement Jan 7, 2024 by Will Larson · across every show →
Everything said about benchmarks, oldest first
Jan 7, 2024 negative
Larson: VC Spending Benchmarks Placate Boards but Do Not Improve Engineering
“You know, this idea that if you just have the right benchmarks, like DCs won't judge you for spending too much in engineering, but it doesn't actually help you get to the right place. It just helps you get your board to be less angry at you.”
Feb 9, 2025 neutral
Nguyen: AI bottleneck is evaluations rather than data as benchmarks saturate
“We are actually getting saturated in all benchmarks. So I think the bottleneck is actually in evaluations that we don't have all the frontier, like evals, like, I don't know GPGA, which is, like, A Google-proof question answering, like, PhD-level intelligence …”
Jul 17, 2025 positive
Shipper: Real-world 'vibe checks' beat standard benchmarks for AI model utility
“I think it's really important to do vibe checks and to call them vibe checks because they're about how does it feel to use this thing and how does it feel to use it for work, for things that you would normally use it for like in your job or in your life. Becau…”
Aug 9, 2025
Turley: Saturated benchmarks mean shipping is the only way to find model failures
“The benchmarks are increasingly saturated. So really you need real world scenarios where your product or model is not actually doing the thing it was supposed to do. And the only way you get that is by shipping because you get back to sort of use case distribu…”
Dec 7, 2025 negative
Chen: Public AI Benchmarks Are Unreliable and Often Contain Wrong Answers
“I don't trust the benchmarks at all. And I think that's for two reasons. So one is, I think a lot of people don't realize, even researchers within the community, they don't realize that the benchmarks themselves are often honestly just wrong. Like they have wr…”
Dec 7, 2025 negative
Chen: Frontier AI Labs Game Benchmarks via Prompt Tweaking and Test Leaks
“Sometimes, yeah, these benchmarks, they accidentally leak in certain ways, or the frontier labs will tweak the way they evaluate their models on these benchmarks. Like they'll tweak their system prompt. Or they'll tweak the number of times they run their model…”
May 24, 2026 neutral
Shipper: AI benchmarks are really one AI-augmented human versus another
“When we are benchmarking against humans, AI against humans, we're actually really always talking about one human using AI versus another human using AI, because AI doesn't use itself. It may be able to in this like slightly somewhat recursive way, but there's …”