benchmarks

5 statements across 3 episodes · 1 bullish · 2 bearish · 3 people on the record · first statement Feb 14, 2025 by Winston Weinberg · across every show →

Everything said about benchmarks, oldest first

Feb 14, 2025 negative
Opinion
Weinberg: Standard AI benchmarks are useless for evaluating legal AI
“Most benchmarks are completely useless for us, right? And so we'll get a model, you know, someone will give us early access to a model and they'll say it's way better on all of these benchmarks and we'll respond. It actually isn't like, it's not used as useful…”
Winston Weinberg Feb 14, 2025 ▶ 9:15 No Priors Ep. 101 | With Harvey CEO and Co-Founder Winston Weinberg
Jan 8, 2026 positive
Assertion Supported
Gil: Chinese open-source AI models rank among highest on benchmarks
“Some of the highest Scoring models against benchmarks now are Chinese models on the open source side. On the closer side, it's still a lot of the US models, but things like Quinn, DeepSeq, et cetera, are doing very well.”
Elad Gil Jan 8, 2026 ▶ 15:14 NVIDIA’s Jensen Huang on Reasoning Models, Robotics, and Refuting the “AI Bubble” Narrative
Jun 26, 2026 negative
Insight
Brown: Scaffolding Easily Inflates AI Benchmark Scores Without Real Gains
“It's really easy to show you can do much better than previous benchmarks or previous, previous models on benchmarks by just, for example, scaffolding a bunch of models together. So if you say, okay, well, we're going to, instead of just running this model once…”
Noam Brown Jun 26, 2026 ▶ 7:03 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
Jun 26, 2026 neutral
Insight
Brown: Benchmark Gains From Routing May Fail in Real-World Use
“One issue you could run into is that you could optimize for certain benchmarks with the routing and then show like, oh yeah, we see this big improvement on these benchmarks. But in real world use cases, it actually ends up not being a significant improvement.”
Noam Brown Jun 26, 2026 ▶ 35:28 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
Jun 26, 2026 neutral
Insight
Brown: Long AI Deliberation Time Is Impractical for Real Workflows
“This idea that the models, you just let them think for a week or whatever, and then they respond, it's, it sounds nice, and yes, the benchmarks look great, but it's not very practical when working because like, okay, you ask the model a question, and then you …”
Noam Brown Jun 26, 2026 ▶ 6:13 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.