Hays: Users have very low trust in AI benchmarks
“I just think people have very, very low trust in benchmarks. At this point, everyone has had enough experience using various models. They have their own sort of internal benchmark.”
Altman: Current AI benchmarks are definitely inadequate
“Definitely not. In some sense, the eval that matters is like, is this being useful to people? You can approximate it by revenue or by amount of usage or like rate of discovery of new knowledge.”
Brown: AI benchmarks must control for test-time compute
“And so I think the proper way to, and so my claim is the proper way to evaluate the models now is you either have some kind of budget for the benchmark, whether it's tokens or cost or time or whatever, or you plot the performance as a function of the amount of…”
Brown: AI Benchmarks in One Year Will Measure Cost and Time per Task
“That's where I think the benchmark thing will be a year from now. It's not going to be how token efficient is it. It's going to be how much money and how much time does it cost to do a specific task?”
Dean: AI benchmarks above 95% accuracy offer diminishing returns due to data leakage
“I think once it hits kind of 95% or something, you get very diminishing returns from really focusing on that benchmark because it's sort of, it's either the case that you've now achieved that capability or there's also the issue of leakage in public data or ve…”
The AI industry is rapidly running out of challenging evaluation benchmarks
“The only thing is we are running out of is really benchmarks. So the improvement on benchmarks, it's kind of like harder to measure.”
Izmailov: AI models can quickly max out defined benchmarks using RL
“And I think we are at the stage where if we define a benchmark and we can make a relevant RL environment, then we can kind of max it out pretty quickly, and so we are going through benchmarks now very, very quickly.”
Frosst: AI benchmark fixation is unhelpful for regulation because benchmarks are easily gamed
“Fixation on particular benchmarks, which can be gamed and can be trained either to do way better on or way worse on. Are not helpful for establishing how the technology can be used and misused.”
Agrawal: Public AI benchmarks are saturated; proprietary evals are required
“The benchmarks are the starting line, but they're by no means the finish line. The benchmarks are kind of saturated, right?... What you actually need is more proprietary evals.”
Kim: Real-world usage will replace saturated benchmarks to measure AI progress
“I feel like we've almost saturated a lot of these evals, and the real, like, metric of, like, how good our models are getting is, I think, gonna be, like, usage, right?”
Mann: New AI benchmarks are fully saturated within 6 to 12 months
“There's this great chart on our world in data that shows that when you release a new benchmark within like six to 12 months, it immediately gets saturated.”
Duffy: AI benchmarks follow a lifecycle from initial idea to saturation
“Essentially there's, I think, a life cycle of a benchmark, right? It starts with an idea, then it gets adopted, and then it gets saturated.”
Bender says most AI benchmarks fail to measure actual capabilities
“Most of the benchmarks that are out there are not reasonable. They lack what's called construct validity, and construct validity is this two-part test of the thing that we are trying to measure is a real thing, and this measurement correlates with it interesti…”
Wang: AI benchmarks are saturated, making model leaders hard to distinguish
“One thing that we see today with the models is that because all the benchmarks that were used today are what's called saturated, i.e., you know, in other words, like all the models do really well at the benchmarks, it's really hard to discern actually which on…”
Yao: Lack of realistic benchmarks is AI's primary bottleneck
“So I think right now the problem is not even that we don't have good methodologies, it's more about we don't have good tasks.”
AI evaluation benchmarks fail for audio, necessitating human aesthetic judgment
“I think in, in all branches of AI, we become slaves to our metrics, and you say, I did this accuracy on this benchmark, and this accuracy on this benchmark, and in the real world, sometimes it doesn't necessarily matter, and these benchmarks are extra terrible…”