AI Model Evaluation

topic on 2 shows · 2 statements across 2 episodes

the MAD Podcast the a16z Podcast

2 statements about AI Model Evaluation, every show

a16z Prediction Not checkable as stated
Angelopoulos: Future AI evaluation will shift to personalized user leaderboards
“Absolutely. Absolutely. It should be personalized just for you. You should understand which models are best for you.”
Anastasios Angelopoulos May 29, 2025 ▶ 26:06 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
MAD Insight
Delangue: Public AI leaderboards are the wrong way to evaluate models
“People are doing evaluation the wrong way. You know, just looking at one massive public leaderboard that doesn't really tell them much about how the model is going to perform on their own use case.”
Clement Delangue Oct 17, 2024 ▶ 1:06:12 The $4.5B Platform Driving the Open Source AI Revolution | Clem Delangue, CEO, Hugging Face

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.