AI evaluation
also referred to as: ai evaluations
4 statements across 4 episodes · 1 bullish · 3 bearish · 4 people on the record · first statement Mar 13, 2025 by Hamel Husain · across every show →
Everything said about AI evaluation, oldest first
Mar 13, 2025 positive
Husain: AI evaluation principles are evergreen unless AGI arrives
“And I found that, like, the subject is pretty evergreen, because we're not, you know, over the last year and a half, like, the same principles apply. And, you know, we're not really talking about Like, you know, using specific tools and APIs is more of a gener…”
Apr 15, 2025 negative
Jan 9, 2026 negative
Hill-Smith: Widely tracked AI benchmarks improve without reflecting general intelligence gains
“Once an eval becomes the thing that everyone's looking at, schools can get better on it without there being a reflection of overall generalized intelligence of these models getting better. That has been true for the last couple of years. It'll be true for the …”
Jun 4, 2026 negative
Petersson: Percentage-based AI benchmarks saturate with noise above 92%
“Even when you're not at a hundred, I think a lot of these evals have a lot of problems in them. So, like, actually, it's, like, if you get to, like, 92 or something like that, many of them, it's, like, then there's, like, there's no, really no difference betwe…”