Assertion Supported AI assessment confidence: 95% certainty 5/5 debate potential 2/5

Angelopoulos: LMArena makes style control the default AI evaluation method

Anastasios Angelopoulos · Beyond Leaderboards: LMArena’s Mission to Make AI Reliable · May 29, 2025 · at 12:55

Anastasios Angelopoulos announces that LMArena is implementing style control by default to isolate substantive model quality from surface-level biases like verbosity.

0:00 / 0:01exact quote · 1.7s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“That's why we're making style control default.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Anastasios Angelopoulos

Opinion
Angelopoulos: Believing chatbot leaderboards are easily gameable is naive
“I have to say, I also just like completely disagree with the foundation of the question. The like implicit assumption is that like chat is easy or that it's even easier than web dev. That's completely false. It's a completely naive perspective that people have…”
Anastasios Angelopoulos May 29, 2025 ▶ 24:57 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Insight
Angelopoulos: Most mission-critical AI queries are subjective, not factual lookups
“In reality, even in such industries, the majority of questions that people ask are subjective. Okay. So the mythology that in hard sciences or in mission critical industries, people just have like cut and dried questions and they just need like a retrieval and…”
Anastasios Angelopoulos May 29, 2025 ▶ 2:50 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Assertion Not checkable as stated
Angelopoulos: Chatbot Arena is immune to model overfitting by design
“Static benchmarks overfit. Why? It's because as Jan said earlier, you're giving the student the same test over and over. You have a model, you test it, you know, you look at whether or not it's improved on a static data set. Then you find another model, you te…”
Anastasios Angelopoulos May 29, 2025 ▶ 21:00 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Prediction Not checkable as stated
Angelopoulos: Future AI evaluation will shift to personalized user leaderboards
“Absolutely. Absolutely. It should be personalized just for you. You should understand which models are best for you.”
Anastasios Angelopoulos May 29, 2025 ▶ 26:06 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Opinion
Angelopoulos: Industry AI evaluation platforms face skepticism over bias
“The fact that we come from Berkeley and from a university really speaks to our scientific approach in neutrality. I think if it came from an industrial lab, people would always have questions about, oh, well, these people are they also training a model and wha…”
Anastasios Angelopoulos May 29, 2025 ▶ 39:36 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Insight
Angelopoulos: AI evaluation performance follows a data scaling law
“Because language models are sort of the intermediary that gets you to this evaluation, there's also a scaling law that comes along with it. Which is to say that the more data you get, the bigger you build the platform, the better you can make your evaluations,…”
Anastasios Angelopoulos May 29, 2025 ▶ 56:18 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.