LM arena

also referred to as: lmarena

16 statements across 1 episodes · 13 bullish · 1 bearish · 4 people on the record · first statement May 29, 2025 by Anastasios Angelopoulos · said 11 times in 1 episodes since 2025 · across every show →

Mentions by year

brought up most by Anastasios Angelopoulos (10), Anjney Midha (1)

tap a year for its mentions
00811512025episodesmentions
0112025episodes it came up in
007.50.51512025episodesmentions per episode
2025 11 mentions in 1 episode

every mention, scene by scene, with the transcript →

Everything said about LM arena, oldest first

May 29, 2025 positive
Disclosure
Angelopoulos: LMArena conducts pre-release model testing for AI developers
“One of the things that we help everybody to do is pre-release testing of their models. Okay. So it's not just that, you know, we work together to evaluate the models are released, but we also try to be their release partners and say, Hey, can we help you guys …”
Anastasios Angelopoulos May 29, 2025 ▶ 4:38 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
May 29, 2025 positive
Assertion Supported
Angelopoulos: Bradley-Terry models converge for AI evaluation, unlike Elo scores
“Okay, let's move from Elo to Bradley Terry because we're actually performing an estimate here instead of just like You know, and the ELO score moves over time. It doesn't converge, but Rally Terry models converge and how do we then construct confidence interva…”
Anastasios Angelopoulos May 29, 2025 ▶ 38:49 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
May 29, 2025 positive
Prediction Not checkable as stated
Angelopoulos: Real-world testing will remain fundamental for evaluating AI agents
“The fundamental is organic, real-world testing with feedback. That's not going to change. I can tell you that that is not going to change.”
Anastasios Angelopoulos May 29, 2025 ▶ 1:44:12 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
May 29, 2025 positive
Assertion Supported
Angelopoulos: LMArena prompt router yields double the performance per dollar
“Now, if you trace the performance, the best performance that, you know, any individual model can give you as part of the router as a function of cost. That's like two X worse than the router. In other words, the router is giving you double the bang for your bu…”
Anastasios Angelopoulos May 29, 2025 ▶ 1:30:12 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
May 29, 2025 positive
Disclosure
Angelopoulos: LMArena is building personalized AI leaderboard tools
“Yeah, and we should be giving you the tools to do that, and we're currently building them.”
Anastasios Angelopoulos May 29, 2025 ▶ 25:56 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
May 29, 2025 negative
Opinion
Angelopoulos: Believing chatbot leaderboards are easily gameable is naive
“I have to say, I also just like completely disagree with the foundation of the question. The like implicit assumption is that like chat is easy or that it's even easier than web dev. That's completely false. It's a completely naive perspective that people have…”
Anastasios Angelopoulos May 29, 2025 ▶ 24:57 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
May 29, 2025 bullish
Prediction Didn’t hold up
Angelopoulos: LMArena will launch Data-Driven Debugging within months
“So we're building a project now that we call data-driven debugging D three. It's, you know, it's a little farther out. It'll come in a couple months.”
Anastasios Angelopoulos May 29, 2025 ▶ 1:26:14 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
May 29, 2025 positive
Prediction Not checkable as stated
Chiang: Personalized leaderboards will improve LMArena data quality
“And in that case, we align the interests of individuals and the platform as a whole, because you don't want to mess up your personal leaderboard. Just like how people these days, when they use social media, They don't like a random post because if they do that…”
Wei-Lin Chiang May 29, 2025 ▶ 1:32:22 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
May 29, 2025 positive
Prediction Held up
Angelopoulos: LMArena will remain open-source as a commercial company
“We're going to keep publishing papers. We're going to keep releasing open source. We're going to keep releasing open data.”
Anastasios Angelopoulos May 29, 2025 ▶ 1:36:21 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
May 29, 2025 positive
Assertion Not checkable as stated
Stoica: Over 70% of daily LMArena prompts are completely unique
“And basically measures out how many more, you know, fresh prompts you have in one day compared to what you've seen in the past three months, right? And by a similarity score of something like 70, 75%, you have over 70 of these prompts are fresh.”
Ion Stoica May 29, 2025 ▶ 20:13 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
May 29, 2025 positive
Assertion Supported
Angelopoulos: LMArena makes style control the default AI evaluation method
“That's why we're making style control default.”
Anastasios Angelopoulos May 29, 2025 ▶ 12:55 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
May 29, 2025
Disclosure
Chiang: LMArena open-sources all code, infrastructure, and models
“All the code infrastructures that we process the data is published as open source and also research blog, paper, and then including prompt leaderboard, we publish the paper. Open source, the models, the code, and everything.”
Wei-Lin Chiang May 29, 2025 ▶ 1:35:01 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
May 29, 2025 positive
Assertion Supported
Angelopoulos: LMArena router model outperforms all constituent models on Chatbot Arena
“When you train a prompt to leaderboard model, which is like, let's say a seven billion parameter model, and then you use it to route on just questions on the arena and everybody's questions, that model does better than any of the constituent models that were u…”
Anastasios Angelopoulos May 29, 2025 ▶ 1:29:10 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
May 29, 2025 positive
Assertion Supported
Midha: LMArena hosts over 280 AI models, up from 12 initially
“In total, that first year there were about 12 models or so, and today it's over 280 or something on the platform.”
Anjney Midha May 29, 2025 ▶ 49:05 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
May 29, 2025 neutral
Assertion Not checkable as stated
Angelopoulos: LMArena measures user preference, not AGI progress
“We don't claim to be an AGI benchmark. We are faithfully representing the preferences of our community.”
Anastasios Angelopoulos May 29, 2025 ▶ 18:08 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
May 29, 2025 positive
Opinion
Angelopoulos: Industry AI evaluation platforms face skepticism over bias
“The fact that we come from Berkeley and from a university really speaks to our scientific approach in neutrality. I think if it came from an industrial lab, people would always have questions about, oh, well, these people are they also training a model and wha…”
Anastasios Angelopoulos May 29, 2025 ▶ 39:36 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.