Staniszewski: Benchmarking text-to-speech models is extremely difficult due to voice differences
“So like even doing benchmarks for text to speech is extremely hard. Because usually different models will have different voices. That already makes them uncomparable.”
Chi: AI data vendors create gimmick benchmarks to sell data
“And actually a lot of that industry has now Built these gimmick style benchmarks as a mechanism to sell their data. And so that, that's become kind of their go-to-market as well.”
Chi: Retiring AI benchmarks is necessary to reflect current real-world knowledge
“There's another component of retiring benchmarks, which I think is, is underappreciated which is that benchmark should also be reflective of the current state of the world.”
Goel: Model Evaluation Is Undervalued; Audio AI Benchmarks Remain Inadequate
“The most undervalued part of this is evaluating your models. Because, like, I think, especially in things like audio. Yeah. There, when we started the benchmarks were pretty non-existent. Even today, I would say the benchmarks are not great.”
Early quantum benchmarks like boson sampling were terribly misleading
“Ok, there's no practical application for those kinds of things. So what you end up getting Are, in some cases, in the early days benchmarks that sounded interesting but were terribly misleading.”
Brown: Long AI Deliberation Time Is Impractical for Real Workflows
“This idea that the models, you just let them think for a week or whatever, and then they respond, it's, it sounds nice, and yes, the benchmarks look great, but it's not very practical when working because like, okay, you ask the model a question, and then you …”
Brown: Scaffolding Easily Inflates AI Benchmark Scores Without Real Gains
“It's really easy to show you can do much better than previous benchmarks or previous, previous models on benchmarks by just, for example, scaffolding a bunch of models together. So if you say, okay, well, we're going to, instead of just running this model once…”
Brown: Benchmark Gains From Routing May Fail in Real-World Use
“One issue you could run into is that you could optimize for certain benchmarks with the routing and then show like, oh yeah, we see this big improvement on these benchmarks. But in real world use cases, it actually ends up not being a significant improvement.”
Shipper: AI benchmarks are really one AI-augmented human versus another
“When we are benchmarking against humans, AI against humans, we're actually really always talking about one human using AI versus another human using AI, because AI doesn't use itself. It may be able to in this like slightly somewhat recursive way, but there's …”
Staniszewski says ElevenLabs speech-to-text models beat industry benchmarks across 100 languages
“Speech to text models that work over a hundred languages and happily beat others on benchmarks all the way through to conversational models of how you loop them together to music, to other domains of audio.”
Friedman: Future SOTA Benchmarks May Come From Swarms of Cheaper AI Models
“It might be that like the next stuff that like is soda on benchmarks is not the most expensive newest foundation model with the most like GPU training. It's like a swarm of lower cost cheaper models working together just like humans do to solve a problem.”
O'Driscoll: a16z wins more top Series A deals than Benchmark
“Andreessen's market share is higher than benchmarks in terms of that. That worked. And I wish Rotman, the guy from DST did. It's higher in terms of the great series A's, right? As a market share, but the hit rate is much lower.”
Gil: Chinese open-source AI models rank among highest on benchmarks
“Some of the highest
Scoring models against benchmarks now are Chinese models on the open source side.
On the closer side, it's still a lot of the US models, but things like Quinn, DeepSeq, et cetera, are doing very well.”
Chen: Public AI Benchmarks Are Unreliable and Often Contain Wrong Answers
“I don't trust the benchmarks at all. And I think that's for two reasons. So one is, I think a lot of people don't realize, even researchers within the community, they don't realize that the benchmarks themselves are often honestly just wrong. Like they have wr…”
Chen: Frontier AI Labs Game Benchmarks via Prompt Tweaking and Test Leaks
“Sometimes, yeah, these benchmarks, they accidentally leak in certain ways, or the frontier labs will tweak the way they evaluate their models on these benchmarks. Like they'll tweak their system prompt. Or they'll tweak the number of times they run their model…”
Osika: AI model benchmarks decay over time due to Goodhart's Law
“I mean, they turn more and more bullshit over time. There's something called good hearts law. So when you start optimizing for a number, that number stops being a good measure for success.”
Turley: Saturated benchmarks mean shipping is the only way to find model failures
“The benchmarks are increasingly saturated. So really you need real world scenarios where your product or model is not actually doing the thing it was supposed to do. And the only way you get that is by shipping because you get back to sort of use case distribu…”
Fortuna: Healthcare AI evaluation must measure worst-case failures over best-of-N
“I think one of the problems is, like, the benchmarks, they always report, like, the best event. So they report, like, you know, what's the best metric they got out of 64 attempts? But in healthcare, we're more interested in, like, what's the worst event, right…”
Shipper: Real-world 'vibe checks' beat standard benchmarks for AI model utility
“I think it's really important to do vibe checks and to call them vibe checks because they're about how does it feel to use this thing and how does it feel to use it for work, for things that you would normally use it for like in your job or in your life. Becau…”
Wu: Reinforcement learning will eventually beat any benchmark with a clear feedback loop
“I think the natural conclusion of RL, which is what we're kind of getting to, is you basically can solve any benchmark, which is insane to think about... Which means like, if you have a clean set of environments, if you have a good feedback loop to decide what…”
Guha: LLMs that fail public benchmarks can excel in specific niches
“Newer LLMs sometimes that don't do so well in benchmarks do much better for your use case.”
Weinberg: Standard AI benchmarks are useless for evaluating legal AI
“Most benchmarks are completely useless for us, right? And so we'll get a model, you know, someone will give us early access to a model and they'll say it's way better on all of these benchmarks and we'll respond. It actually isn't like, it's not used as useful…”
Nguyen: AI bottleneck is evaluations rather than data as benchmarks saturate
“We are actually getting saturated in all benchmarks. So I think the bottleneck is actually in evaluations that we don't have all the frontier, like evals, like, I don't know GPGA, which is, like, A Google-proof question answering, like, PhD-level intelligence …”
GitHub CPO: Every Existing Public AI Benchmark Can Be Gamed
“I don't like benchmarks out there by the way, because you could game every single one of them. In my opinion, but what I do like about benchmarking is that it gives you a view into a set of scenarios that then you could then figure out, are you getting better …”
Swyx: Frontier AI labs distinguish themselves by adopting new benchmarks
“The labs that are not that frontier will keep measuring themselves on last year's benchmarks. And then the labs that are actually frontier will tell you about benchmarks you've never heard of.”
Bank: Donor-Backed Institutions Have Much Lower 'Embarrassment Risk' Tolerance
“And the last one, which I think is the most delicate, is variance risk, or what I'll call with clients embarrassment risk, which is how far behind benchmarks, peers, whomever, are you willing to be at any given time?
That one is something that is generally unk…”
Carlini: Users should build personalized AI benchmarks instead of relying on public leaderboards
“The argument that I tried to lay out in this post is that more people should make benchmarks that are tailored to them.”
Narayanan: AI developers over-optimize models for benchmarks over real-world performance
“When there is so much pressure to do well on these benchmarks, developers are intentionally or unintentionally optimizing these models In ways that look good on the benchmarks, but don't look good in real world evaluation.”
Albrecht: Imbue reproduced 500-1,000 examples per dataset to stop eval contamination
“Let's just reproduce, you know, 500 to a thousand examples for every single one of these data sets ourselves and just make sure that this data is definitely not in the, you know, the training set. So we did that and then we're able to like now be confident abo…”
Conover: AI model developers are absolutely overfitting to public evaluation benchmarks
“And I think the work around over, you know, overfitting on the test, I think is like that. 100% is happening.”
Shulman: AI evaluation benchmarks are far worse in audio than text
“As flawed as these benchmarks are in text, they're way worse in audio.”
Larson: VC Spending Benchmarks Placate Boards but Do Not Improve Engineering
“You know, this idea that if you just have the right benchmarks, like DCs won't judge you for spending too much in engineering, but it doesn't actually help you get to the right place. It just helps you get your board to be less angry at you.”
Wood: ARK portfolios have under 5% overlap with major equity benchmarks
“Less than 10% of our portfolio is, or I should say less than, there's much less than a 10% overlap between us and any, you know, it's usually less than five percent.”
Kim: Top VCs maximize ownership and squeeze LPs out of Series A
“If there is, at the earliest stages, a company that is High quality with high quality investors coming in. LPs are probably the last in line in terms of getting access to that. So imagine a seed funded company by one of our fund managers, a high quality firm l…”