Benchmarks

topic on 14 shows · 34 statements across 30 episodes

the Y Combinator Startup Podcast Cheeky Pint We Live to Build American Optimist David Senra Latent Space Lenny's Podcast the Neon Show No Priors the Official SaaStr Podcast Capital Allocators Catalyst the a16z Podcast 20VC

34 statements about Benchmarks, every show

Staniszewski: Benchmarking text-to-speech models is extremely difficult due to voice differences
“So like even doing benchmarks for text to speech is extremely hard. Because usually different models will have different voices. That already makes them uncomparable.”
Mati Staniszewski Sep 9, 2026 ▶ 55:09 Building One of AI’s Fastest-Growing Companies | Mati Staniszewski, ElevenLabs
a16z Opinion
Chi: AI data vendors create gimmick benchmarks to sell data
“And actually a lot of that industry has now Built these gimmick style benchmarks as a mechanism to sell their data. And so that, that's become kind of their go-to-market as well.”
Ryan Chi Sep 9, 2026 ▶ 9:15 Inside the Race to Measure Frontier Intelligence
a16z Insight
Chi: Retiring AI benchmarks is necessary to reflect current real-world knowledge
“There's another component of retiring benchmarks, which I think is, is underappreciated which is that benchmark should also be reflective of the current state of the world.”
Ryan Chi Sep 9, 2026 ▶ 12:49 Inside the Race to Measure Frontier Intelligence
NEON SHOW Opinion
Goel: Model Evaluation Is Undervalued; Audio AI Benchmarks Remain Inadequate
“The most undervalued part of this is evaluating your models. Because, like, I think, especially in things like audio. Yeah. There, when we started the benchmarks were pretty non-existent. Even today, I would say the benchmarks are not great.”
Karan Goel Aug 7, 2026 ▶ 16:32 The Billion Dollar AI Lab Founder Who Sees The Future First | Karan Goel, Founder & CEO of Cartesia
CATALYST Opinion
Early quantum benchmarks like boson sampling were terribly misleading
“Ok, there's no practical application for those kinds of things. So what you end up getting Are, in some cases, in the early days benchmarks that sounded interesting but were terribly misleading.”
Bob Sorenson Jul 16, 2026 ▶ 9:31 When will quantum computing have its breakout moment?
NO PRIORS Insight
Brown: Long AI Deliberation Time Is Impractical for Real Workflows
“This idea that the models, you just let them think for a week or whatever, and then they respond, it's, it sounds nice, and yes, the benchmarks look great, but it's not very practical when working because like, okay, you ask the model a question, and then you …”
Noam Brown Jun 26, 2026 ▶ 6:13 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
NO PRIORS Insight
Brown: Scaffolding Easily Inflates AI Benchmark Scores Without Real Gains
“It's really easy to show you can do much better than previous benchmarks or previous, previous models on benchmarks by just, for example, scaffolding a bunch of models together. So if you say, okay, well, we're going to, instead of just running this model once…”
Noam Brown Jun 26, 2026 ▶ 7:03 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
NO PRIORS Insight
Brown: Benchmark Gains From Routing May Fail in Real-World Use
“One issue you could run into is that you could optimize for certain benchmarks with the routing and then show like, oh yeah, we see this big improvement on these benchmarks. But in real world use cases, it actually ends up not being a significant improvement.”
Noam Brown Jun 26, 2026 ▶ 35:28 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
Shipper: AI benchmarks are really one AI-augmented human versus another
“When we are benchmarking against humans, AI against humans, we're actually really always talking about one human using AI versus another human using AI, because AI doesn't use itself. It may be able to in this like slightly somewhat recursive way, but there's …”
Dan Shipper May 24, 2026 ▶ 48:07 AI predictions: Job markets, Codex beats Claude, and the death of org charts | Dan Shipper
CHEEKY PINT Assertion Partly supported
Staniszewski says ElevenLabs speech-to-text models beat industry benchmarks across 100 languages
“Speech to text models that work over a hundred languages and happily beat others on benchmarks all the way through to conversational models of how you loop them together to music, to other domains of audio.”
Mati Staniszewski Apr 14, 2026 ▶ 9:43 The world of voice AI, with Mati Staniszewski of ElevenLabs
Y COMBINATOR Prediction Not checkable as stated
Friedman: Future SOTA Benchmarks May Come From Swarms of Cheaper AI Models
“It might be that like the next stuff that like is soda on benchmarks is not the most expensive newest foundation model with the most like GPU training. It's like a swarm of lower cost cheaper models working together just like humans do to solve a problem.”
Jared Friedman Feb 21, 2026 ▶ 17:03 The AI Agent Economy Is Here · Y Combinator
20VC Assertion Supported
O'Driscoll: a16z wins more top Series A deals than Benchmark
“Andreessen's market share is higher than benchmarks in terms of that. That worked. And I wish Rotman, the guy from DST did. It's higher in terms of the great series A's, right? As a market share, but the hit rate is much lower.”
Rory O'Driscoll Jan 15, 2026 ▶ 40:16 Anthropic’s $10B Raise | a16z’s $15B Fund: Is the Middle Dead in VC? | How OpenAI Could Go to Zero? · 20VC with Harry Stebbings
NO PRIORS Assertion Supported
Gil: Chinese open-source AI models rank among highest on benchmarks
“Some of the highest Scoring models against benchmarks now are Chinese models on the open source side. On the closer side, it's still a lot of the US models, but things like Quinn, DeepSeq, et cetera, are doing very well.”
Elad Gil Jan 8, 2026 ▶ 15:14 NVIDIA’s Jensen Huang on Reasoning Models, Robotics, and Refuting the “AI Bubble” Narrative
Chen: Public AI Benchmarks Are Unreliable and Often Contain Wrong Answers
“I don't trust the benchmarks at all. And I think that's for two reasons. So one is, I think a lot of people don't realize, even researchers within the community, they don't realize that the benchmarks themselves are often honestly just wrong. Like they have wr…”
Edwin Chen Dec 7, 2025 ▶ 18:01 The $1B Al company training ChatGPT, Claude & Gemini on the path to responsible AGI | Edwin Chen
LENNY'S PODCAST Assertion Not checkable as stated
Chen: Frontier AI Labs Game Benchmarks via Prompt Tweaking and Test Leaks
“Sometimes, yeah, these benchmarks, they accidentally leak in certain ways, or the frontier labs will tweak the way they evaluate their models on these benchmarks. Like they'll tweak their system prompt. Or they'll tweak the number of times they run their model…”
Edwin Chen Dec 7, 2025 ▶ 19:32 The $1B Al company training ChatGPT, Claude & Gemini on the path to responsible AGI | Edwin Chen
20VC Insight
Osika: AI model benchmarks decay over time due to Goodhart's Law
“I mean, they turn more and more bullshit over time. There's something called good hearts law. So when you start optimizing for a number, that number stops being a good measure for success.”
Anton Osika Aug 18, 2025 ▶ 1:02:27 Lovable CEO, Anton Osika: The State of Foundation Models, Grok vs OpenAI, and Replit vs Bolt · 20VC with Harry Stebbings
Turley: Saturated benchmarks mean shipping is the only way to find model failures
“The benchmarks are increasingly saturated. So really you need real world scenarios where your product or model is not actually doing the thing it was supposed to do. And the only way you get that is by shipping because you get back to sort of use case distribu…”
Nick Turley Aug 9, 2025 ▶ 1:13:57 Inside ChatGPT: The fastest growing product in history | Nick Turley (OpenAI)
Fortuna: Healthcare AI evaluation must measure worst-case failures over best-of-N
“I think one of the problems is, like, the benchmarks, they always report, like, the best event. So they report, like, you know, what's the best metric they got out of 64 attempts? But in healthcare, we're more interested in, like, what's the worst event, right…”
Brendan Fortuna Jul 29, 2025 ▶ 19:09 ⚡️Using RFT to Build Clinical Superintelligence
Shipper: Real-world 'vibe checks' beat standard benchmarks for AI model utility
“I think it's really important to do vibe checks and to call them vibe checks because they're about how does it feel to use this thing and how does it feel to use it for work, for things that you would normally use it for like in your job or in your life. Becau…”
Dan Shipper Jul 17, 2025 ▶ 27:24 The AI-native startup: 5 products, 7-figure revenue, 100% AI-written code. | Dan Shipper (Every)
AMERICAN OPTIMIST Prediction Not checkable as stated
Wu: Reinforcement learning will eventually beat any benchmark with a clear feedback loop
“I think the natural conclusion of RL, which is what we're kind of getting to, is you basically can solve any benchmark, which is insane to think about... Which means like, if you have a clean set of environments, if you have a good feedback loop to decide what…”
Scott Wu Jun 12, 2025 ▶ 15:09 From Math Prodigy to AI Genius: How Scott Wu Built Devin · Joe Lonsdale
Guha: LLMs that fail public benchmarks can excel in specific niches
“Newer LLMs sometimes that don't do so well in benchmarks do much better for your use case.”
Rabbi Guha Mar 4, 2025 ▶ 21:39 Founder Market Fit Got Them $4.2M Before a Product Existed
NO PRIORS Opinion
Weinberg: Standard AI benchmarks are useless for evaluating legal AI
“Most benchmarks are completely useless for us, right? And so we'll get a model, you know, someone will give us early access to a model and they'll say it's way better on all of these benchmarks and we'll respond. It actually isn't like, it's not used as useful…”
Winston Weinberg Feb 14, 2025 ▶ 9:15 No Priors Ep. 101 | With Harvey CEO and Co-Founder Winston Weinberg
Nguyen: AI bottleneck is evaluations rather than data as benchmarks saturate
“We are actually getting saturated in all benchmarks. So I think the bottleneck is actually in evaluations that we don't have all the frontier, like evals, like, I don't know GPGA, which is, like, A Google-proof question answering, like, PhD-level intelligence …”
Karina Nguyen Feb 9, 2025 ▶ 10:49 OpenAI researcher on why soft skills are the future of work | Karina Nguyen
SAASTR Opinion
GitHub CPO: Every Existing Public AI Benchmark Can Be Gamed
“I don't like benchmarks out there by the way, because you could game every single one of them. In my opinion, but what I do like about benchmarking is that it gives you a view into a set of scenarios that then you could then figure out, are you getting better …”
Mario Rodriguez Jan 10, 2025 ▶ 28:06 Adding AI to SaaS: Inside the AI Product Strategies of Figma, Cloudflare, GitHub and Ramp
Swyx: Frontier AI labs distinguish themselves by adopting new benchmarks
“The labs that are not that frontier will keep measuring themselves on last year's benchmarks. And then the labs that are actually frontier will tell you about benchmarks you've never heard of.”
Shawn Wang Jan 1, 2025 ▶ 1:08:52 2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
Bank: Donor-Backed Institutions Have Much Lower 'Embarrassment Risk' Tolerance
“And the last one, which I think is the most delicate, is variance risk, or what I'll call with clients embarrassment risk, which is how far behind benchmarks, peers, whomever, are you willing to be at any given time? That one is something that is generally unk…”
Matt Bank Nov 25, 2024 ▶ 13:58 Matt Bank - "GEMs" of Risk, Asset Allocation, and Manager Selection (EP.419)
Carlini: Users should build personalized AI benchmarks instead of relying on public leaderboards
“The argument that I tried to lay out in this post is that more people should make benchmarks that are tailored to them.”
Nicholas Carlini Aug 28, 2024 ▶ 39:10 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
20VC Assertion Not checkable as stated
Narayanan: AI developers over-optimize models for benchmarks over real-world performance
“When there is so much pressure to do well on these benchmarks, developers are intentionally or unintentionally optimizing these models In ways that look good on the benchmarks, but don't look good in real world evaluation.”
Arvind Narayanan Aug 28, 2024 ▶ 18:25 Arvind Narayanan: AI Scaling Myths, The Core Bottlenecks in AI Today & The Future of Models | E1195 · 20VC with Harry Stebbings
LATENT SPACE Disclosure
Albrecht: Imbue reproduced 500-1,000 examples per dataset to stop eval contamination
“Let's just reproduce, you know, 500 to a thousand examples for every single one of these data sets ourselves and just make sure that this data is definitely not in the, you know, the training set. So we did that and then we're able to like now be confident abo…”
Josh Albrecht Jun 25, 2024 ▶ 1:00:11 State of the Art: Training 70B LLMs on 10,000 H100 clusters
LATENT SPACE Assertion Not checkable as stated
Conover: AI model developers are absolutely overfitting to public evaluation benchmarks
“And I think the work around over, you know, overfitting on the test, I think is like that. 100% is happening.”
Mike Conover Jun 11, 2024 ▶ 58:21 How AI is Eating Finance - with Mike Conover of Brightwave
Shulman: AI evaluation benchmarks are far worse in audio than text
“As flawed as these benchmarks are in text, they're way worse in audio.”
Mikey Shulman Mar 14, 2024 ▶ 55:42 Making Transformers Sing - with Mikey Shulman of Suno
Larson: VC Spending Benchmarks Placate Boards but Do Not Improve Engineering
“You know, this idea that if you just have the right benchmarks, like DCs won't judge you for spending too much in engineering, but it doesn't actually help you get to the right place. It just helps you get your board to be less angry at you.”
Will Larson Jan 7, 2024 ▶ 49:55 The engineering mindset | Will Larson (Carta, Stripe, Uber, Calm, Digg)
20VC Disclosure
Wood: ARK portfolios have under 5% overlap with major equity benchmarks
“Less than 10% of our portfolio is, or I should say less than, there's much less than a 10% overlap between us and any, you know, it's usually less than five percent.”
Cathie Wood Nov 14, 2022 ▶ 20:20 Cathie Wood: Elon & Twitter; Why Facebook is a Value Stock Now; ARK's Performance | 20VC #949 · 20VC with Harry Stebbings
20VC Insight
Kim: Top VCs maximize ownership and squeeze LPs out of Series A
“If there is, at the earliest stages, a company that is High quality with high quality investors coming in. LPs are probably the last in line in terms of getting access to that. So imagine a seed funded company by one of our fund managers, a high quality firm l…”
Michael Kim Jul 6, 2016 ▶ 11:59 20VC: What It Takes To Raise A VC Fund & Investing in First Time Fund Managers with Michael Kim @ Cendana Capital

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.