Shawn Wang: Fixed-capability LLM inference costs drop 100x to 1000x annually
“In language models, it is roughly 100 to a thousand times every 12 to 18 months for the same given level of LMSYS ELO.”
Angelopoulos: Arena Dropped 'LM' to Broaden Beyond Language Models
“So, so we wanted to maybe broaden a little bit. And we were the first Serena, so we feel like let's kind of try to own that.”
Wu: OpenAI was first to launch stateful responses API
“Obviously we were the first one to launch responses API, but like a couple of other people have kind of adopted, I think Grok has it in their API. I think I saw LMSYS just did something”
Lambert: LMSYS is probably setting up a deep research arena
“I mean, they're probably setting up a deep research arena, because that's the data that, I mean, if I was open AI working on deep research, that's the data that I want, and there are competitors, and LMSYS is the entity that has the market placement to set it …”
Lambert: Meta withholding its leading benchmark model is bad execution
“But to be a model that claims to be open and then not release the model that is your leading claim is just, like, that is, like, bad execution.”
Chen: Researchers Degrade AI Factuality Just to Boost LMSYS Rankings
“A lot of researchers, they'll tell us that their VPs make them focus on increasing their rank on LMSYS. And so I've had researchers explicitly tell me that they're okay with making their models worse. Add factuality works at following instructions as long as i…”
Beauchamp: 5,000 completions needed to accurately ELO rank community models
“We think it takes about 5000 completions to get an accurate signal.”
Angelopoulos: Static benchmarks are intrinsically unable to evaluate generative models
“Static benchmarks are intrinsically, to some extent, unable to measure generative model performance. And the reason is because you cannot Pre-annotate all the outputs of a generative model. You change the model. It's like the distribution of your data is chang…”
Chiang: Coding questions drive 20% to 30% of Chatbot Arena usage
“We do see a lot of like developers come to the site asking polling questions. Only 30%.”
Chiang: Chatbot Arena almost died after launch due to low engagement
“At some point, almost died. Because as you can imagine, this leaderboard depends on user, like part of, like community engagement participation. If no one comes to vote, Tomorrow then no deal.”
Angelopoulos: LMSYS controls for markdown and lists in Arena rankings
“We have, you know, five, six different style components that have to do with markdown headers and bulleted lists and so on that we add here.”
Angelopoulos: LMSYS considers default style control but avoids imposing opinions
“We consider that we're still actively considering it. It's just, you know, once you make that step, once you take that step, you're introducing your opinion. And I'm not, you know, why should our opinion be the one? That's kind of a community choice. We could …”
Angelopoulos: LMSYS wants to integrate live code execution in Chatbot Arena
“For example, it'd be great if we could execute code within Arena. It'd be fantastic. We want to do it.”
Angelopoulos: OpenAI o1 crushed Chatbot Arena, proving the benchmark isn't saturated
“So there's this model and it crushed the benchmark. You know, it's just like really like a big gap. And what that's telling us is that it's not saturated yet. And so it's still measuring some signal that was encouraging point.”
Angelopoulos: Five-model selection bias is tiny compared to voter variability
“We don't do that right now, partially because we kind of have know from simulations that the amount of selection bias you incur with these five things is just not huge. It's not huge in comparison to the variability that you get from the, from just regular hum…”
Angelopoulos: Live voter data asymptotically eliminates pre-release ELO bias
“What happened is that over time, because we're getting new data, it'll get adjusted down. So if there's any bias that gets introduced at that stage in the long run, it actually doesn't matter because asymptotically, basically like in the long run, there's way …”
Chiang: There are currently no good benchmarks for evaluating LLM routers
“Right now, currently, there seems to be the, one of the end point when we developed this project was like, there's just no good benchmark for a router.”
Angelopoulos: Highly effective LLM routers only need simple heuristics like length
“Well, I think that you can build a very, very simple router that is very effective. So let me give you an example. You can build a great router with one parameter, and the parameter is just like, I'm gonna check if my question is hard, and if it's hard, then I…”
Angelopoulos: Chatbot Arena is decoupling from LMSYS as co-creators shift focus
“Sort of Chatbot Arena has, of course, like, kind of become its own thing, and Lianmin and Ying, who are, you know, created LMSYS, have kind of, like, moved on to working on SGLang, and now They're doing other projects that are sort of originating from LMSS. An…”
Yi Tay: Distilled open-source model variants disappeared after failing to climb LMSYS
“When people realize that, like, this, like, turning on the GPT-IV tab and running some DPO is not going to give them the reward signal that they want anymore, right? Then all these variants gone, right? You know, there was this era where there's, wow, there's …”
Lambert: Chatbot Arena is the best available evaluation benchmark for LLMs
“I have, if we make it to evaluation, I'd pretty much say that Chat Arena is the best limited evaluation that people have to learn how to use language models, and like, It's very valuable data”