AI Benchmarks

topic on 11 shows · 16 statements across 16 episodes

Latent Space Lenny's Podcast No Priors the Startup Ideas Podcast Invest Like the Best Sourcery the MAD Podcast the a16z Podcast Big Technology TBPN 20VC

16 statements about AI Benchmarks, every show

TBPN Opinion
Hays: Users have very low trust in AI benchmarks
“I just think people have very, very low trust in benchmarks. At this point, everyone has had enough experience using various models. They have their own sort of internal benchmark.”
Jordi Hays Sep 4, 2026 ▶ 3:34 Model Mayhem, GPT-6 Astra, Why Nvidia Bought Hugging Face | Diet TBPN
Altman: Current AI benchmarks are definitely inadequate
“Definitely not. In some sense, the eval that matters is like, is this being useful to people? You can approximate it by revenue or by amount of usage or like rate of discovery of new knowledge.”
Sam Altman Jul 28, 2026 ▶ 28:27 Sam Altman on AGI, Compute, and Human Agency · Invest Like The Best
NO PRIORS Insight
Brown: AI benchmarks must control for test-time compute
“And so I think the proper way to, and so my claim is the proper way to evaluate the models now is you either have some kind of budget for the benchmark, whether it's tokens or cost or time or whatever, or you plot the performance as a function of the amount of…”
Noam Brown Jun 26, 2026 ▶ 4:01 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
STARTUP IDEAS Prediction Not checkable as stated
Brown: AI Benchmarks in One Year Will Measure Cost and Time per Task
“That's where I think the benchmark thing will be a year from now. It's not going to be how token efficient is it. It's going to be how much money and how much time does it cost to do a specific task?”
Riley Brown Apr 27, 2026 ▶ 46:29 Stop using Claude. Start using Codex?
Dean: AI benchmarks above 95% accuracy offer diminishing returns due to data leakage
“I think once it hits kind of 95% or something, you get very diminishing returns from really focusing on that benchmark because it's sort of, it's either the case that you've now achieved that capability or there's also the issue of leakage in public data or ve…”
Jeff Dean Feb 12, 2026 ▶ 12:01 The AI Frontier: from Gemini 3 Deep Think distilling to Flash — Jeff Dean
MAD Assertion Not checkable as stated
The AI industry is rapidly running out of challenging evaluation benchmarks
“The only thing is we are running out of is really benchmarks. So the improvement on benchmarks, it's kind of like harder to measure.”
Sebastian Raschka Jan 29, 2026 ▶ 37:56 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
MAD Insight
Izmailov: AI models can quickly max out defined benchmarks using RL
“And I think we are at the stage where if we define a benchmark and we can make a relevant RL environment, then we can kind of max it out pretty quickly, and so we are going through benchmarks now very, very quickly.”
Pavel Izmailov Jan 15, 2026 ▶ 26:42 The Evaluators Are Being Evaluated — Pavel Izmailov (Anthropic/NYU)
20VC Insight
Frosst: AI benchmark fixation is unhelpful for regulation because benchmarks are easily gamed
“Fixation on particular benchmarks, which can be gamed and can be trained either to do way better on or way worse on. Are not helpful for establishing how the technology can be used and misused.”
Nick Frosst Sep 1, 2025 ▶ 1:08:39 Cohere Founder, Nick Frosst: How To Compete with OpenAI & Anthropic, and Sam Altman’s AI Disservice · 20VC with Harry Stebbings
SOURCERY Insight
Agrawal: Public AI benchmarks are saturated; proprietary evals are required
“The benchmarks are the starting line, but they're by no means the finish line. The benchmarks are kind of saturated, right?... What you actually need is more proprietary evals.”
Apoorv Agrawal Aug 27, 2025 ▶ 28:46 Understanding OpenAI’s $500B Valuation · Sourcery with Molly O'Shea
a16z Insight
Kim: Real-world usage will replace saturated benchmarks to measure AI progress
“I feel like we've almost saturated a lot of these evals, and the real, like, metric of, like, how good our models are getting is, I think, gonna be, like, usage, right?”
Christina Kim Aug 8, 2025 ▶ 9:15 GPT-5 and Agents Breakdown – w/ OpenAI Researchers Isa Fulford & Christina Kim
LENNY'S PODCAST Assertion Supported
Mann: New AI benchmarks are fully saturated within 6 to 12 months
“There's this great chart on our world in data that shows that when you release a new benchmark within like six to 12 months, it immediately gets saturated.”
Ben Mann Jul 20, 2025 ▶ 10:27 Anthropic co-founder: AGI predictions, leaving OpenAI, what keeps him up at night | Ben Mann
Duffy: AI benchmarks follow a lifecycle from initial idea to saturation
“Essentially there's, I think, a life cycle of a benchmark, right? It starts with an idea, then it gets adopted, and then it gets saturated.”
Alex Duffy Jun 11, 2025 ▶ 21:11 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
Bender says most AI benchmarks fail to measure actual capabilities
“Most of the benchmarks that are out there are not reasonable. They lack what's called construct validity, and construct validity is this two-part test of the thing that we are trying to measure is a real thing, and this measurement correlates with it interesti…”
Emily M. Bender May 14, 2025 ▶ 14:30 AI’s Drawbacks: Environmental Damage, Bad Benchmarks, Outsourcing Thinking
BIG TECHNOLOGY Assertion Not checkable as stated
Wang: AI benchmarks are saturated, making model leaders hard to distinguish
“One thing that we see today with the models is that because all the benchmarks that were used today are what's called saturated, i.e., you know, in other words, like all the models do really well at the benchmarks, it's really hard to discern actually which on…”
Alexandr Wang Dec 11, 2024 ▶ 53:56 AI Predictions for 2025: Geopolitics, Agents, and Data Scaling — With Alexandr Wang
Yao: Lack of realistic benchmarks is AI's primary bottleneck
“So I think right now the problem is not even that we don't have good methodologies, it's more about we don't have good tasks.”
Shunyu Yao Sep 27, 2024 ▶ 31:08 Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
NO PRIORS Insight
AI evaluation benchmarks fail for audio, necessitating human aesthetic judgment
“I think in, in all branches of AI, we become slaves to our metrics, and you say, I did this accuracy on this benchmark, and this accuracy on this benchmark, and in the real world, sometimes it doesn't necessarily matter, and these benchmarks are extra terrible…”
Mikey Shulman May 16, 2024 ▶ 7:11 No Priors Ep. 64 | With Suno CEO and Co-Founder Mikey Shulman

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.