benchmarks
5 statements across 3 episodes · 1 bullish · 2 bearish · 3 people on the record · first statement Feb 14, 2025 by Winston Weinberg · across every show →
Everything said about benchmarks, oldest first
Feb 14, 2025 negative
Weinberg: Standard AI benchmarks are useless for evaluating legal AI
“Most benchmarks are completely useless for us, right? And so we'll get a model, you know, someone will give us early access to a model and they'll say it's way better on all of these benchmarks and we'll respond. It actually isn't like, it's not used as useful…”
Jan 8, 2026 positive
Jun 26, 2026 negative
Brown: Scaffolding Easily Inflates AI Benchmark Scores Without Real Gains
“It's really easy to show you can do much better than previous benchmarks or previous, previous models on benchmarks by just, for example, scaffolding a bunch of models together. So if you say, okay, well, we're going to, instead of just running this model once…”
Jun 26, 2026 neutral
Brown: Benchmark Gains From Routing May Fail in Real-World Use
“One issue you could run into is that you could optimize for certain benchmarks with the routing and then show like, oh yeah, we see this big improvement on these benchmarks. But in real world use cases, it actually ends up not being a significant improvement.”
Jun 26, 2026 neutral
Brown: Long AI Deliberation Time Is Impractical for Real Workflows
“This idea that the models, you just let them think for a week or whatever, and then they respond, it's, it sounds nice, and yes, the benchmarks look great, but it's not very practical when working because like, okay, you ask the model a question, and then you …”