AI evaluation

2 statements across 2 episodes · 1 bullish · 0 bearish · 2 people on the record · first statement May 29, 2025 by Anastasios Angelopoulos · across every show →

Everything said about AI evaluation, oldest first

May 29, 2025 positive
Insight
Angelopoulos: AI evaluation performance follows a data scaling law
“Because language models are sort of the intermediary that gets you to this evaluation, there's also a scaling law that comes along with it. Which is to say that the more data you get, the bigger you build the platform, the better you can make your evaluations,…”
Anastasios Angelopoulos May 29, 2025 ▶ 56:18 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Sep 9, 2026 neutral
Prediction Not checkable as stated
Chi: Legible enterprise evals will be AI adoption's biggest long-term bottleneck
“And I think long-term that will be actually the biggest bottleneck, our ability to take companies and their evals and make them legible because that's how we'll figure out what signal we hill climb on and where we actually adopt.”
Ryan Chi Sep 9, 2026 ▶ 6:41 Inside the Race to Measure Frontier Intelligence
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.