Assertion Not checkable as stated
Arena's anonymous Nano Banana test moved Google's stock and product roadmap
“I mean, that moment alone changed Google's like roadmap. Market share. Seriously. I mean, Google stock, billions of dollars are moving because of Nano.”
Insight
Angelopoulos: Static benchmarks are intrinsically unable to evaluate generative models
“Static benchmarks are intrinsically, to some extent, unable to measure generative model performance. And the reason is because you cannot Pre-annotate all the outputs of a generative model. You change the model. It's like the distribution of your data is chang…”
Opinion
Arena's organic user prompts provide realism that Artificial Analysis lacks
“They have arenas, but the arenas are not based on organic usage. Like the thing that distinguishes our platform versus theirs is that the users are actually inputting their own use case. They're actually asking their own question. And that gives a level of rea…”
Assertion Supported
Arena sampled open-source models at 60/40, debunking Leaderboard Illusion paper
“But, you know, there, for example said that we were, that we only sampled, like, nine percent open source models and, like, you know, 60%, like, closed source models, and this created a gap between open and closed source. But in reality, we're actually really …”
Disclosure
Arena's public leaderboard will never adopt Gartner-style pay-to-play models
“You can't pay to get on the public leaderboard. It's not like a Gartner in that sense. It's not like any of these, like you know, pay to play systems, never going to be like that. Models are going to be listed on the leaderboard, whether or not the providers p…”
Opinion
Angelopoulos rejects claims that Cognition's Devin is dead
“Devin's not gone. Devin's everywhere.”
Assertion Supported
Angelopoulos: OpenAI o1 crushed Chatbot Arena, proving the benchmark isn't saturated
“So there's this model and it crushed the benchmark. You know, it's just like really like a big gap. And what that's telling us is that it's not saturated yet. And so it's still measuring some signal that was encouraging point.”
Assertion Supported
Angelopoulos: The Chatbot Arena leaderboard is currently not an apples-to-apples comparison
“None of the leaderboard currently is apples to apples, because you have, like, Gemini Flash, you have, you know, all sorts of tiny models, like Llama Like, eight B and four or five B are not apples to apples.”
Insight
Angelopoulos: Highly effective LLM routers only need simple heuristics like length
“Well, I think that you can build a very, very simple router that is very effective. So let me give you an example. You can build a great router with one parameter, and the parameter is just like, I'm gonna check if my question is hard, and if it's hard, then I…”
Disclosure
Angelopoulos: LMArena receives only standard enterprise inference discounts
“No, no, we get discounts, but they're, but they are standard enterprise discounts. The same that would be given to any other customer.”
Assertion Not checkable as stated
Arena processes tens of millions of conversations monthly, totaling 250 million
“We have probably two hundred and fifty million conversations that happen over the course of the platform. We're on the order of, you know, mid tens of millions of conversations every month that are happening on the platform.”
Prediction Not checkable as stated
Angelopoulos: Academic paper figures will soon be generated by AI models
“Soon we're not going to be even making them for our papers. We're, they're just going to be, our paper figures are going to be made by Emily.”
Assertion Not checkable as stated
Angelopoulos: LMArena has released more real-world AI data than almost anyone
“We've probably released more data than basically anybody on the real world use cases of AI.”
Disclosure
Angelopoulos: LMSYS considers default style control but avoids imposing opinions
“We consider that we're still actively considering it. It's just, you know, once you make that step, once you take that step, you're introducing your opinion. And I'm not, you know, why should our opinion be the one? That's kind of a community choice. We could …”
Assertion Not checkable as stated
Angelopoulos: Five-model selection bias is tiny compared to voter variability
“We don't do that right now, partially because we kind of have know from simulations that the amount of selection bias you incur with these five things is just not huge. It's not huge in comparison to the variability that you get from the, from just regular hum…”
Insight
Angelopoulos: Live voter data asymptotically eliminates pre-release ELO bias
“What happened is that over time, because we're getting new data, it'll get adjusted down. So if there's any bias that gets introduced at that stage in the long run, it actually doesn't matter because asymptotically, basically like in the long run, there's way …”
Disclosure
Angelopoulos: Arena Dropped 'LM' to Broaden Beyond Language Models
“So, so we wanted to maybe broaden a little bit. And we were the first Serena, so we feel like let's kind of try to own that.”
Assertion Supported
Angelopoulos: Arena received grants from Sequoia and a16z before incorporating
“He was not, you know, A-sixteen was not the only one to do this. We also had a great grant from Sequoia, but Ansh was in particular quite, quite supportive of us and, you know, gave us some resources in order to continue building out Arena before we even We're…”
Assertion Not checkable as stated
LMArena funds all model inference running on its platform
“We fund all of the inference on the platform.”
Assertion Not checkable as stated
Angelopoulos: 25% of LMArena platform users write software for a living
“25% of the people on our platform, for example, do software for a living.”
Assertion Not checkable as stated
Angelopoulos: About half of LMArena users are authenticated
“About half of our users now are login.”
Assertion Supported
Gradio scaled Arena to 1 million monthly active users before migration
“Gradio scaled us to a million Mal.”
Disclosure
Angelopoulos: LMArena's top expenses are free-tier inference, hiring, and SF office
“Primarily inference that funds the free usage of the platform and then also hiring, of course, headcount. We have an office, you know. That's an SF.”
Disclosure
Angelopoulos: LMArena to launch video evaluations by early next year
“Video we're soon to launch on the site at some point, you know, later this year or early next.”