benchmarks

7 statements across 6 episodes · 1 bullish · 3 bearish · 5 people on the record · first statement Jan 7, 2024 by Will Larson · across every show →

Everything said about benchmarks, oldest first

Jan 7, 2024 negative
Insight
Larson: VC Spending Benchmarks Placate Boards but Do Not Improve Engineering
“You know, this idea that if you just have the right benchmarks, like DCs won't judge you for spending too much in engineering, but it doesn't actually help you get to the right place. It just helps you get your board to be less angry at you.”
Will Larson Jan 7, 2024 ▶ 49:55 The engineering mindset | Will Larson (Carta, Stripe, Uber, Calm, Digg)
Feb 9, 2025 neutral
Opinion
Nguyen: AI bottleneck is evaluations rather than data as benchmarks saturate
“We are actually getting saturated in all benchmarks. So I think the bottleneck is actually in evaluations that we don't have all the frontier, like evals, like, I don't know GPGA, which is, like, A Google-proof question answering, like, PhD-level intelligence …”
Karina Nguyen Feb 9, 2025 ▶ 10:49 OpenAI researcher on why soft skills are the future of work | Karina Nguyen
Jul 17, 2025 positive
Insight
Shipper: Real-world 'vibe checks' beat standard benchmarks for AI model utility
“I think it's really important to do vibe checks and to call them vibe checks because they're about how does it feel to use this thing and how does it feel to use it for work, for things that you would normally use it for like in your job or in your life. Becau…”
Dan Shipper Jul 17, 2025 ▶ 27:24 The AI-native startup: 5 products, 7-figure revenue, 100% AI-written code. | Dan Shipper (Every)
Aug 9, 2025
Insight
Turley: Saturated benchmarks mean shipping is the only way to find model failures
“The benchmarks are increasingly saturated. So really you need real world scenarios where your product or model is not actually doing the thing it was supposed to do. And the only way you get that is by shipping because you get back to sort of use case distribu…”
Nick Turley Aug 9, 2025 ▶ 1:13:57 Inside ChatGPT: The fastest growing product in history | Nick Turley (OpenAI)
Dec 7, 2025 negative
Opinion
Chen: Public AI Benchmarks Are Unreliable and Often Contain Wrong Answers
“I don't trust the benchmarks at all. And I think that's for two reasons. So one is, I think a lot of people don't realize, even researchers within the community, they don't realize that the benchmarks themselves are often honestly just wrong. Like they have wr…”
Edwin Chen Dec 7, 2025 ▶ 18:01 The $1B Al company training ChatGPT, Claude & Gemini on the path to responsible AGI | Edwin Chen
Dec 7, 2025 negative
Assertion Not checkable as stated
Chen: Frontier AI Labs Game Benchmarks via Prompt Tweaking and Test Leaks
“Sometimes, yeah, these benchmarks, they accidentally leak in certain ways, or the frontier labs will tweak the way they evaluate their models on these benchmarks. Like they'll tweak their system prompt. Or they'll tweak the number of times they run their model…”
Edwin Chen Dec 7, 2025 ▶ 19:32 The $1B Al company training ChatGPT, Claude & Gemini on the path to responsible AGI | Edwin Chen
May 24, 2026 neutral
Insight
Shipper: AI benchmarks are really one AI-augmented human versus another
“When we are benchmarking against humans, AI against humans, we're actually really always talking about one human using AI versus another human using AI, because AI doesn't use itself. It may be able to in this like slightly somewhat recursive way, but there's …”
Dan Shipper May 24, 2026 ▶ 48:07 AI predictions: Job markets, Codex beats Claude, and the death of org charts | Dan Shipper
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.