Opinion
Angelopoulos: Believing chatbot leaderboards are easily gameable is naive
“I have to say, I also just like completely disagree with the foundation of the question. The like implicit assumption is that like chat is easy or that it's even easier than web dev. That's completely false. It's a completely naive perspective that people have…”
Insight
Angelopoulos: Most mission-critical AI queries are subjective, not factual lookups
“In reality, even in such industries, the majority of questions that people ask are subjective. Okay. So the mythology that in hard sciences or in mission critical industries, people just have like cut and dried questions and they just need like a retrieval and…”
Assertion Not checkable as stated
Angelopoulos: Chatbot Arena is immune to model overfitting by design
“Static benchmarks overfit. Why? It's because as Jan said earlier, you're giving the student the same test over and over. You have a model, you test it, you know, you look at whether or not it's improved on a static data set. Then you find another model, you te…”
Assertion Supported
Angelopoulos: LMArena makes style control the default AI evaluation method
“That's why we're making style control default.”
Prediction Not checkable as stated
Angelopoulos: Future AI evaluation will shift to personalized user leaderboards
“Absolutely. Absolutely. It should be personalized just for you. You should understand which models are best for you.”
Opinion
Angelopoulos: Industry AI evaluation platforms face skepticism over bias
“The fact that we come from Berkeley and from a university really speaks to our scientific approach in neutrality. I think if it came from an industrial lab, people would always have questions about, oh, well, these people are they also training a model and wha…”
Insight
Angelopoulos: AI evaluation performance follows a data scaling law
“Because language models are sort of the intermediary that gets you to this evaluation, there's also a scaling law that comes along with it. Which is to say that the more data you get, the bigger you build the platform, the better you can make your evaluations,…”
Insight
Angelopoulos: AI leaderboards can utilize any form of interaction feedback
“Pairwise comparison feedback is not the only kind of feedback that we can use to construct leaderboards. We can construct leaderboards with any form of feedback.”
Assertion Supported
Angelopoulos: LMArena router model outperforms all constituent models on Chatbot Arena
“When you train a prompt to leaderboard model, which is like, let's say a seven billion parameter model, and then you use it to route on just questions on the arena and everybody's questions, that model does better than any of the constituent models that were u…”
Assertion Supported
Angelopoulos: LMArena prompt router yields double the performance per dollar
“Now, if you trace the performance, the best performance that, you know, any individual model can give you as part of the router as a function of cost. That's like two X worse than the router. In other words, the router is giving you double the bang for your bu…”
Insight
Angelopoulos: Top AI researchers avoid companies building purely proprietary technology
“The best people don't want to hole up at a company and develop a bunch of proprietary technology that, you know, is never going to be released.”
Insight
Angelopoulos: High refusal rates do not make AI models inherently superior
“It's not necessarily the model that's like most, like refuses the most to answer these like queries that people ask necessarily better. Some people want a model that's more controllable. Some people want a model that's going to say whatever they want. Some peo…”
Assertion Open · timeframe May 2026
Angelopoulos: Human evaluators prefer longer AI responses given equal content
“It's true that people vote for longer responses, you know, preferentially over shorter responses, even given the same contents or well-known human bias.”
Assertion Not checkable as stated
Angelopoulos: LMArena measures user preference, not AGI progress
“We don't claim to be an AGI benchmark. We are faithfully representing the preferences of our community.”
Disclosure
Angelopoulos: LMArena is building personalized AI leaderboard tools
“Yeah, and we should be giving you the tools to do that, and we're currently building them.”
Assertion Supported
Angelopoulos: Bradley-Terry models converge for AI evaluation, unlike Elo scores
“Okay, let's move from Elo to Bradley Terry because we're actually performing an estimate here instead of just like You know, and the ELO score moves over time. It doesn't converge, but Rally Terry models converge and how do we then construct confidence interva…”
Insight
Angelopoulos: Reinforcement learning allows AI models to surpass human teachers
“And supervised learning, you can only do as well as the best human that you have. Because what's happening is that you're learning from the teacher. In reinforcement learning, you're learning from the world. You're able to learn things better than the best hum…”
Assertion Not checkable as stated
Angelopoulos: Chatbot Arena has 1M+ monthly users and 150M+ conversations
“A lot of people don't know this, but ShopBot Arena is Used by like a million plus monthly users. We get like, you know, tens of thousands of votes on a daily basis. We have like over like, you know, a hundred fifty million conversations that have been had on t…”
Prediction Held up
Angelopoulos: LMArena will remain open-source as a commercial company
“We're going to keep publishing papers. We're going to keep releasing open source. We're going to keep releasing open data.”
Prediction Not checkable as stated
Angelopoulos: Real-world testing will remain fundamental for evaluating AI agents
“The fundamental is organic, real-world testing with feedback. That's not going to change. I can tell you that that is not going to change.”
Disclosure
Angelopoulos: LMArena conducts pre-release model testing for AI developers
“One of the things that we help everybody to do is pre-release testing of their models. Okay. So it's not just that, you know, we work together to evaluate the models are released, but we also try to be their release partners and say, Hey, can we help you guys …”
Prediction Didn’t hold up
Angelopoulos: LMArena will launch Data-Driven Debugging within months
“So we're building a project now that we call data-driven debugging D three. It's, you know, it's a little farther out. It'll come in a couple months.”