AI benchmarks

3 statements across 3 episodes · 0 bullish · 1 bearish · 3 people on the record · first statement Sep 27, 2024 by Shunyu Yao · across every show →

Everything said about AI benchmarks, oldest first

Sep 27, 2024 negative
Insight
Yao: Lack of realistic benchmarks is AI's primary bottleneck
“So I think right now the problem is not even that we don't have good methodologies, it's more about we don't have good tasks.”
Shunyu Yao Sep 27, 2024 ▶ 31:08 Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
Jun 11, 2025 neutral
Insight
Duffy: AI benchmarks follow a lifecycle from initial idea to saturation
“Essentially there's, I think, a life cycle of a benchmark, right? It starts with an idea, then it gets adopted, and then it gets saturated.”
Alex Duffy Jun 11, 2025 ▶ 21:11 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
Feb 12, 2026
Insight
Dean: AI benchmarks above 95% accuracy offer diminishing returns due to data leakage
“I think once it hits kind of 95% or something, you get very diminishing returns from really focusing on that benchmark because it's sort of, it's either the case that you've now achieved that capability or there's also the issue of leakage in public data or ve…”
Jeff Dean Feb 12, 2026 ▶ 12:01 The AI Frontier: from Gemini 3 Deep Think distilling to Flash — Jeff Dean
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.