Anastasios Angelopoulos, co-founder of LMArena, responds to host Anjney Midha regarding how AI evaluation must adapt as models evolve into autonomous agents performing long-horizon tasks.
“The fundamental is organic, real-world testing with feedback. That's not going to change. I can tell you that that is not going to change.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Anastasios Angelopoulos
Opinion
Angelopoulos: Believing chatbot leaderboards are easily gameable is naive
“I have to say, I also just like completely disagree with the foundation of the question. The like implicit assumption is that like chat is easy or that it's even easier than web dev. That's completely false. It's a completely naive perspective that people have…”
Anastasios AngelopoulosMay 29, 2025▶ 24:57Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Insight
Angelopoulos: Most mission-critical AI queries are subjective, not factual lookups
“In reality, even in such industries, the majority of questions that people ask are subjective. Okay. So the mythology that in hard sciences or in mission critical industries, people just have like cut and dried questions and they just need like a retrieval and…”
Anastasios AngelopoulosMay 29, 2025▶ 2:50Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
AssertionNot checkable as stated
Angelopoulos: Chatbot Arena is immune to model overfitting by design
“Static benchmarks overfit. Why? It's because as Jan said earlier, you're giving the student the same test over and over. You have a model, you test it, you know, you look at whether or not it's improved on a static data set. Then you find another model, you te…”
Anastasios AngelopoulosMay 29, 2025▶ 21:00Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
AssertionSupported
Angelopoulos: LMArena makes style control the default AI evaluation method
“That's why we're making style control default.”
Anastasios AngelopoulosMay 29, 2025▶ 12:55Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
PredictionNot checkable as stated
Angelopoulos: Future AI evaluation will shift to personalized user leaderboards
“Absolutely. Absolutely. It should be personalized just for you. You should understand which models are best for you.”
Anastasios AngelopoulosMay 29, 2025▶ 26:06Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Opinion
Angelopoulos: Industry AI evaluation platforms face skepticism over bias
“The fact that we come from Berkeley and from a university really speaks to our scientific approach in neutrality. I think if it came from an industrial lab, people would always have questions about, oh, well, these people are they also training a model and wha…”
Anastasios AngelopoulosMay 29, 2025▶ 39:36Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 1,000 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.