Nov 1, 2024 · 41m · latent-space

In the Arena: How LMSys changed LLM Benchmarking Forever

Anastasios Angelopoulos · 17m spoken Wei-Lin Chiang · 10m spoken Shawn Wang · 5m spoken Alessio Fanelli · 3m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the Latent Space podcast, LMSYS researchers Wei-Lin Chiang and Anastasios Angelopoulos discuss the origins, statistical methodology, and expansion of Chatbot Arena into the premier open-source benchmark for large language models.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 23.8% of the talking time here. How this is scored →

The hosts as informed peer 4.7 Guest teaching 4.1 Guest disagreement 0.9 The hosts pushing back 1.8
05100:0015:0030:000:04–5:08 · The hosts as informed peer 3/10 Introductions and Academic Research Backgrounds Hosts open with a lighthearted intro and ask about the origins of Chatbot Arena and Vicuna. The guests provide a collaborative narrative on early fine-tuning experiments following Stanford's Alpaca.5:08–9:02 · The hosts as informed peer 4/10 Dynamic Human Evaluation Versus Static Benchmarks Alessio asks why LMSYS chose a human voting arena over traditional static benchmarks. The guests explain the fundamental mathematical and practical limits of static benchmarks for open-ended generative models.9:03–12:57 · The hosts as informed peer 3/10 Early Community Adoption, Growth Trajectory, and Trust Swyx asks about early growth mechanics and user distribution skews. Wei-Lin and Anastasios candidly discuss how the platform almost died until model provider rivalry created organic viewer demand.12:59–18:10 · The hosts as informed peer 7/10 Controlling for Human Biases and Style in ELO Ratings The conversation turns to statistical control for length and formatting biases in Bradley-Terry ratings. Anastasios explains logistic regression confounder adjustment, which Swyx matches with his background in econometrics and quantitative trading.18:13–20:58 · The hosts as informed peer 5/10 Expanding Categories: Coding, Math, and Hard Prompts Swyx lightly challenges LMSYS for not making style control the default setting, noting that neutrality is itself an opinion. The guests explain their conservative philosophy on imposing subjective priors.20:59–25:50 · The hosts as informed peer 6/10 Red Teaming, Skill Disparities, and System Security Alessio probes how Arena handles high-skill domains like red teaming and system-level data exfiltration. The guests describe framing jailbreaks as explicit human-versus-model competitive games.25:51–28:34 · The hosts as informed peer 5/10 The Impact of OpenAI o1 on LLM Benchmarks Swyx brings up academic critiques that OpenAI o1 is not an apples-to-apples comparison due to test-time search compute. Anastasios counters that leaderboards already compare heterogeneous model sizes and latencies.28:36–34:26 · The hosts as informed peer 6/10 Private Model Testing, Selection Bias, and ELO Stability Swyx asks directly about community controversy regarding frontier labs testing multiple private models to cherry-pick peaks. Anastasios provides a rigorous statistical defence using Bonferroni corrections and asymptotic data updates.34:26–37:55 · The hosts as informed peer 5/10 RouteLLM and the Role of Model Routers Swyx asks whether routing models must be as intelligent as frontier models to route effectively. Anastasios reframes routing as an easy single-parameter cost/performance threshold rather than complex reasoning.37:56–40:50 · The hosts as informed peer 3/10 LMSYS Evolution, Arena Independence, and Community Call for Contributions Alessio asks about LMSYS project decoupling and graduation. The guests explain the organizational changes and invite community pull requests for interactive REPL code execution.0:04–5:08 · Guest teaching 2/10 Introductions and Academic Research Backgrounds Hosts open with a lighthearted intro and ask about the origins of Chatbot Arena and Vicuna. The guests provide a collaborative narrative on early fine-tuning experiments following Stanford's Alpaca.5:08–9:02 · Guest teaching 5/10 Dynamic Human Evaluation Versus Static Benchmarks Alessio asks why LMSYS chose a human voting arena over traditional static benchmarks. The guests explain the fundamental mathematical and practical limits of static benchmarks for open-ended generative models.9:03–12:57 · Guest teaching 2/10 Early Community Adoption, Growth Trajectory, and Trust Swyx asks about early growth mechanics and user distribution skews. Wei-Lin and Anastasios candidly discuss how the platform almost died until model provider rivalry created organic viewer demand.12:59–18:10 · Guest teaching 6/10 Controlling for Human Biases and Style in ELO Ratings The conversation turns to statistical control for length and formatting biases in Bradley-Terry ratings. Anastasios explains logistic regression confounder adjustment, which Swyx matches with his background in econometrics and quantitative trading.18:13–20:58 · Guest teaching 3/10 Expanding Categories: Coding, Math, and Hard Prompts Swyx lightly challenges LMSYS for not making style control the default setting, noting that neutrality is itself an opinion. The guests explain their conservative philosophy on imposing subjective priors.20:59–25:50 · Guest teaching 4/10 Red Teaming, Skill Disparities, and System Security Alessio probes how Arena handles high-skill domains like red teaming and system-level data exfiltration. The guests describe framing jailbreaks as explicit human-versus-model competitive games.25:51–28:34 · Guest teaching 5/10 The Impact of OpenAI o1 on LLM Benchmarks Swyx brings up academic critiques that OpenAI o1 is not an apples-to-apples comparison due to test-time search compute. Anastasios counters that leaderboards already compare heterogeneous model sizes and latencies.28:36–34:26 · Guest teaching 7/10 Private Model Testing, Selection Bias, and ELO Stability Swyx asks directly about community controversy regarding frontier labs testing multiple private models to cherry-pick peaks. Anastasios provides a rigorous statistical defence using Bonferroni corrections and asymptotic data updates.34:26–37:55 · Guest teaching 5/10 RouteLLM and the Role of Model Routers Swyx asks whether routing models must be as intelligent as frontier models to route effectively. Anastasios reframes routing as an easy single-parameter cost/performance threshold rather than complex reasoning.37:56–40:50 · Guest teaching 2/10 LMSYS Evolution, Arena Independence, and Community Call for Contributions Alessio asks about LMSYS project decoupling and graduation. The guests explain the organizational changes and invite community pull requests for interactive REPL code execution.0:04–5:08 · Guest disagreement 0/10 Introductions and Academic Research Backgrounds Hosts open with a lighthearted intro and ask about the origins of Chatbot Arena and Vicuna. The guests provide a collaborative narrative on early fine-tuning experiments following Stanford's Alpaca.5:08–9:02 · Guest disagreement 1/10 Dynamic Human Evaluation Versus Static Benchmarks Alessio asks why LMSYS chose a human voting arena over traditional static benchmarks. The guests explain the fundamental mathematical and practical limits of static benchmarks for open-ended generative models.9:03–12:57 · Guest disagreement 0/10 Early Community Adoption, Growth Trajectory, and Trust Swyx asks about early growth mechanics and user distribution skews. Wei-Lin and Anastasios candidly discuss how the platform almost died until model provider rivalry created organic viewer demand.12:59–18:10 · Guest disagreement 0/10 Controlling for Human Biases and Style in ELO Ratings The conversation turns to statistical control for length and formatting biases in Bradley-Terry ratings. Anastasios explains logistic regression confounder adjustment, which Swyx matches with his background in econometrics and quantitative trading.18:13–20:58 · Guest disagreement 1/10 Expanding Categories: Coding, Math, and Hard Prompts Swyx lightly challenges LMSYS for not making style control the default setting, noting that neutrality is itself an opinion. The guests explain their conservative philosophy on imposing subjective priors.20:59–25:50 · Guest disagreement 0/10 Red Teaming, Skill Disparities, and System Security Alessio probes how Arena handles high-skill domains like red teaming and system-level data exfiltration. The guests describe framing jailbreaks as explicit human-versus-model competitive games.25:51–28:34 · Guest disagreement 3/10 The Impact of OpenAI o1 on LLM Benchmarks Swyx brings up academic critiques that OpenAI o1 is not an apples-to-apples comparison due to test-time search compute. Anastasios counters that leaderboards already compare heterogeneous model sizes and latencies.28:36–34:26 · Guest disagreement 2/10 Private Model Testing, Selection Bias, and ELO Stability Swyx asks directly about community controversy regarding frontier labs testing multiple private models to cherry-pick peaks. Anastasios provides a rigorous statistical defence using Bonferroni corrections and asymptotic data updates.34:26–37:55 · Guest disagreement 2/10 RouteLLM and the Role of Model Routers Swyx asks whether routing models must be as intelligent as frontier models to route effectively. Anastasios reframes routing as an easy single-parameter cost/performance threshold rather than complex reasoning.37:56–40:50 · Guest disagreement 0/10 LMSYS Evolution, Arena Independence, and Community Call for Contributions Alessio asks about LMSYS project decoupling and graduation. The guests explain the organizational changes and invite community pull requests for interactive REPL code execution.0:04–5:08 · The hosts pushing back 0/10 Introductions and Academic Research Backgrounds Hosts open with a lighthearted intro and ask about the origins of Chatbot Arena and Vicuna. The guests provide a collaborative narrative on early fine-tuning experiments following Stanford's Alpaca.5:08–9:02 · The hosts pushing back 2/10 Dynamic Human Evaluation Versus Static Benchmarks Alessio asks why LMSYS chose a human voting arena over traditional static benchmarks. The guests explain the fundamental mathematical and practical limits of static benchmarks for open-ended generative models.9:03–12:57 · The hosts pushing back 1/10 Early Community Adoption, Growth Trajectory, and Trust Swyx asks about early growth mechanics and user distribution skews. Wei-Lin and Anastasios candidly discuss how the platform almost died until model provider rivalry created organic viewer demand.12:59–18:10 · The hosts pushing back 1/10 Controlling for Human Biases and Style in ELO Ratings The conversation turns to statistical control for length and formatting biases in Bradley-Terry ratings. Anastasios explains logistic regression confounder adjustment, which Swyx matches with his background in econometrics and quantitative trading.18:13–20:58 · The hosts pushing back 3/10 Expanding Categories: Coding, Math, and Hard Prompts Swyx lightly challenges LMSYS for not making style control the default setting, noting that neutrality is itself an opinion. The guests explain their conservative philosophy on imposing subjective priors.20:59–25:50 · The hosts pushing back 2/10 Red Teaming, Skill Disparities, and System Security Alessio probes how Arena handles high-skill domains like red teaming and system-level data exfiltration. The guests describe framing jailbreaks as explicit human-versus-model competitive games.25:51–28:34 · The hosts pushing back 3/10 The Impact of OpenAI o1 on LLM Benchmarks Swyx brings up academic critiques that OpenAI o1 is not an apples-to-apples comparison due to test-time search compute. Anastasios counters that leaderboards already compare heterogeneous model sizes and latencies.28:36–34:26 · The hosts pushing back 4/10 Private Model Testing, Selection Bias, and ELO Stability Swyx asks directly about community controversy regarding frontier labs testing multiple private models to cherry-pick peaks. Anastasios provides a rigorous statistical defence using Bonferroni corrections and asymptotic data updates.34:26–37:55 · The hosts pushing back 2/10 RouteLLM and the Role of Model Routers Swyx asks whether routing models must be as intelligent as frontier models to route effectively. Anastasios reframes routing as an easy single-parameter cost/performance threshold rather than complex reasoning.37:56–40:50 · The hosts pushing back 0/10 LMSYS Evolution, Arena Independence, and Community Call for Contributions Alessio asks about LMSYS project decoupling and graduation. The guests explain the organizational changes and invite community pull requests for interactive REPL code execution.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 46.1% · guest 53.9%0:00 · the hosts 46.1% · guest 53.9%3:00 · the hosts 18% · guest 82%3:00 · the hosts 18% · guest 82%6:00 · the hosts 6.1% · guest 93.9%6:00 · the hosts 6.1% · guest 93.9%9:00 · the hosts 20.7% · guest 79.3%9:00 · the hosts 20.7% · guest 79.3%12:00 · the hosts 20.2% · guest 79.8%12:00 · the hosts 20.2% · guest 79.8%15:00 · the hosts 12.2% · guest 87.8%15:00 · the hosts 12.2% · guest 87.8%18:00 · the hosts 21.6% · guest 78.4%18:00 · the hosts 21.6% · guest 78.4%21:00 · the hosts 43.6% · guest 56.4%21:00 · the hosts 43.6% · guest 56.4%24:00 · the hosts 18.8% · guest 81.2%24:00 · the hosts 18.8% · guest 81.2%27:00 · the hosts 32.6% · guest 67.4%27:00 · the hosts 32.6% · guest 67.4%30:00 · the hosts 10.5% · guest 89.5%30:00 · the hosts 10.5% · guest 89.5%33:00 · the hosts 35.2% · guest 64.8%33:00 · the hosts 35.2% · guest 64.8%36:00 · the hosts 21.5% · guest 78.5%36:00 · the hosts 21.5% · guest 78.5%39:00 · the hosts 29.4% · guest 70.6%39:00 · the hosts 29.4% · guest 70.6%
Sharpest disagreement ▶ 28:03 Dismissing the apples-to-apples benchmark criticism

Anastasios firmly counters academic complaints regarding o1's inference compute advantage by pointing out that comparing 8B to 405B models is already heterogeneous.

Hardest push from the hosts ▶ 29:00 Confronting private lab testing suspicions

Swyx directly raises widespread community suspicion that top AI labs game ELO rankings by A/B testing multiple candidate checkpoints and releasing only the highest score.

Biggest teaching moment ▶ 31:16 Statistical breakdown of winner's curse and selection bias

Anastasios educates the hosts on how Bonferroni corrections and asymptotic data collection prevent selection bias from distorting live leaderboard ratings.

The host holds their own ▶ 17:20 Connecting Arena de-biasing to quantitative econometric trading

Swyx demonstrates his technical depth by showing how LMSYS's regression-based style de-biasing maps directly to causal confounder control in financial modeling.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Introductions and Academic Research Backgrounds 3200 Hosts open with a lighthearted intro and ask about the origins of Chatbot Arena and Vicuna. The guests provide a collaborative narrative on early fine-tuning experiments following Stanford's Alpaca.
Dynamic Human Evaluation Versus Static Benchmarks 4512 Alessio asks why LMSYS chose a human voting arena over traditional static benchmarks. The guests explain the fundamental mathematical and practical limits of static benchmarks for open-ended generative models.
Early Community Adoption, Growth Trajectory, and Trust 3201 Swyx asks about early growth mechanics and user distribution skews. Wei-Lin and Anastasios candidly discuss how the platform almost died until model provider rivalry created organic viewer demand.
Controlling for Human Biases and Style in ELO Ratings 7601 The conversation turns to statistical control for length and formatting biases in Bradley-Terry ratings. Anastasios explains logistic regression confounder adjustment, which Swyx matches with his background in econometrics and quantitative trading.
Expanding Categories: Coding, Math, and Hard Prompts 5313 Swyx lightly challenges LMSYS for not making style control the default setting, noting that neutrality is itself an opinion. The guests explain their conservative philosophy on imposing subjective priors.
Red Teaming, Skill Disparities, and System Security 6402 Alessio probes how Arena handles high-skill domains like red teaming and system-level data exfiltration. The guests describe framing jailbreaks as explicit human-versus-model competitive games.
The Impact of OpenAI o1 on LLM Benchmarks 5533 Swyx brings up academic critiques that OpenAI o1 is not an apples-to-apples comparison due to test-time search compute. Anastasios counters that leaderboards already compare heterogeneous model sizes and latencies.
Private Model Testing, Selection Bias, and ELO Stability 6724 Swyx asks directly about community controversy regarding frontier labs testing multiple private models to cherry-pick peaks. Anastasios provides a rigorous statistical defence using Bonferroni corrections and asymptotic data updates.
RouteLLM and the Role of Model Routers 5522 Swyx asks whether routing models must be as intelligent as frontier models to route effectively. Anastasios reframes routing as an easy single-parameter cost/performance threshold rather than complex reasoning.
LMSYS Evolution, Arena Independence, and Community Call for Contributions 3200 Alessio asks about LMSYS project decoupling and graduation. The guests explain the organizational changes and invite community pull requests for interactive REPL code execution.

Statements from this episode (17)

Assertion Not checkable as stated
Chiang: Vicuna Showed Open-Weight Models Could Match ChatGPT Quality
“And people were very excited about it because it kind of like demonstrate open way model can reach this conversation capability similar to ChatGPT.”
Wei-Lin Chiang Nov 1, 2024 ▶ 3:01
Insight
Angelopoulos: Static benchmarks are intrinsically unable to evaluate generative models
“Static benchmarks are intrinsically, to some extent, unable to measure generative model performance. And the reason is because you cannot Pre-annotate all the outputs of a generative model. You change the model. It's like the distribution of your data is chang…”
Anastasios Angelopoulos Nov 1, 2024 ▶ 6:40
Insight
Chiang: Online dynamic benchmarks are slower and more expensive than offline benchmarks
“This kind of like online dynamic benchmark is slow, is more expensive than Standing benchmark, offline benchmark, where people still need it, like when they build models, they need static benchmark to track.”
Wei-Lin Chiang Nov 1, 2024 ▶ 7:48
Assertion Not checkable as stated
Chiang: Coding questions drive 20% to 30% of Chatbot Arena usage
“We do see a lot of like developers come to the site asking polling questions. Only 30%.”
Wei-Lin Chiang Nov 1, 2024 ▶ 10:36
Assertion Not checkable as stated
Chiang: Chatbot Arena almost died after launch due to low engagement
“At some point, almost died. Because as you can imagine, this leaderboard depends on user, like part of, like community engagement participation. If no one comes to vote, Tomorrow then no deal.”
Wei-Lin Chiang Nov 1, 2024 ▶ 11:10
Assertion Supported
Swyx: Humans demonstrably prefer longer outputs in LLM evaluation
“The classic one for human preference evaluation is humans demonstrably prefer longer contexts or longer outputs, which is actually something that we don't necessarily want.”
Shawn Wang Nov 1, 2024 ▶ 13:05
Assertion Supported
Angelopoulos: Chatbot Arena scores are calculated via logistic regression
“The arena score that we show on our leaderboard is a particular type of linear model, right? It's a linear model that takes, it's a logistic regression that takes model identities and fits them against human preference, right? So it regresses human preference …”
Anastasios Angelopoulos Nov 1, 2024 ▶ 14:57
Assertion Partly supported
Angelopoulos: LMSYS controls for markdown and lists in Arena rankings
“We have, you know, five, six different style components that have to do with markdown headers and bulleted lists and so on that we add here.”
Anastasios Angelopoulos Nov 1, 2024 ▶ 16:24
Disclosure
Angelopoulos: LMSYS considers default style control but avoids imposing opinions
“We consider that we're still actively considering it. It's just, you know, once you make that step, once you take that step, you're introducing your opinion. And I'm not, you know, why should our opinion be the one? That's kind of a community choice. We could …”
Anastasios Angelopoulos Nov 1, 2024 ▶ 18:27
Disclosure
Angelopoulos: LMSYS wants to integrate live code execution in Chatbot Arena
“For example, it'd be great if we could execute code within Arena. It'd be fantastic. We want to do it.”
Anastasios Angelopoulos Nov 1, 2024 ▶ 21:49
Assertion Supported
Angelopoulos: OpenAI o1 crushed Chatbot Arena, proving the benchmark isn't saturated
“So there's this model and it crushed the benchmark. You know, it's just like really like a big gap. And what that's telling us is that it's not saturated yet. And so it's still measuring some signal that was encouraging point.”
Anastasios Angelopoulos Nov 1, 2024 ▶ 27:20
Assertion Supported
Angelopoulos: The Chatbot Arena leaderboard is currently not an apples-to-apples comparison
“None of the leaderboard currently is apples to apples, because you have, like, Gemini Flash, you have, you know, all sorts of tiny models, like Llama Like, eight B and four or five B are not apples to apples.”
Anastasios Angelopoulos Nov 1, 2024 ▶ 28:03
Assertion Not checkable as stated
Angelopoulos: Five-model selection bias is tiny compared to voter variability
“We don't do that right now, partially because we kind of have know from simulations that the amount of selection bias you incur with these five things is just not huge. It's not huge in comparison to the variability that you get from the, from just regular hum…”
Anastasios Angelopoulos Nov 1, 2024 ▶ 31:47
Insight
Angelopoulos: Live voter data asymptotically eliminates pre-release ELO bias
“What happened is that over time, because we're getting new data, it'll get adjusted down. So if there's any bias that gets introduced at that stage in the long run, it actually doesn't matter because asymptotically, basically like in the long run, there's way …”
Anastasios Angelopoulos Nov 1, 2024 ▶ 32:16
Assertion Not checkable as stated
Chiang: There are currently no good benchmarks for evaluating LLM routers
“Right now, currently, there seems to be the, one of the end point when we developed this project was like, there's just no good benchmark for a router.”
Wei-Lin Chiang Nov 1, 2024 ▶ 35:57
Insight
Angelopoulos: Highly effective LLM routers only need simple heuristics like length
“Well, I think that you can build a very, very simple router that is very effective. So let me give you an example. You can build a great router with one parameter, and the parameter is just like, I'm gonna check if my question is hard, and if it's hard, then I…”
Anastasios Angelopoulos Nov 1, 2024 ▶ 36:33
Disclosure
Angelopoulos: Chatbot Arena is decoupling from LMSYS as co-creators shift focus
“Sort of Chatbot Arena has, of course, like, kind of become its own thing, and Lianmin and Ying, who are, you know, created LMSYS, have kind of, like, moved on to working on SGLang, and now They're doing other projects that are sort of originating from LMSS. An…”
Anastasios Angelopoulos Nov 1, 2024 ▶ 38:23
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.