Chatbot Arena

11 statements across 2 episodes · 3 bullish · 0 bearish · 3 people on the record · first statement Nov 1, 2024 by Wei-Lin Chiang · said 25 times in 5 episodes since 2024 · across every show →

Mentions by year

brought up most by Shawn Wang (9), Alessio Fanelli (6), Anastasios Angelopoulos (4), Wei-Lin Chiang (3)

tap a year for its mentions
0013225320242025episodesmentions
02320242025episodes it came up in
007.51.515320242025episodesmentions per episode
2025 4 mentions in 3 episodes 1 per episode
2024 21 mentions in 2 episodes 11 per episode

every mention, scene by scene, with the transcript →

Everything said about Chatbot Arena, oldest first

Nov 1, 2024 neutral
Insight
Chiang: Online dynamic benchmarks are slower and more expensive than offline benchmarks
“This kind of like online dynamic benchmark is slow, is more expensive than Standing benchmark, offline benchmark, where people still need it, like when they build models, they need static benchmark to track.”
Wei-Lin Chiang Nov 1, 2024 ▶ 7:48 In the Arena: How LMSys changed LLM Benchmarking Forever
Nov 1, 2024 positive
Disclosure
Angelopoulos: LMSYS wants to integrate live code execution in Chatbot Arena
“For example, it'd be great if we could execute code within Arena. It'd be fantastic. We want to do it.”
Anastasios Angelopoulos Nov 1, 2024 ▶ 21:49 In the Arena: How LMSys changed LLM Benchmarking Forever
Nov 1, 2024 neutral
Assertion Not checkable as stated
Chiang: Coding questions drive 20% to 30% of Chatbot Arena usage
“We do see a lot of like developers come to the site asking polling questions. Only 30%.”
Wei-Lin Chiang Nov 1, 2024 ▶ 10:36 In the Arena: How LMSys changed LLM Benchmarking Forever
Nov 1, 2024 neutral
Disclosure
Angelopoulos: Chatbot Arena is decoupling from LMSYS as co-creators shift focus
“Sort of Chatbot Arena has, of course, like, kind of become its own thing, and Lianmin and Ying, who are, you know, created LMSYS, have kind of, like, moved on to working on SGLang, and now They're doing other projects that are sort of originating from LMSS. An…”
Anastasios Angelopoulos Nov 1, 2024 ▶ 38:23 In the Arena: How LMSys changed LLM Benchmarking Forever
Nov 1, 2024 neutral
Assertion Not checkable as stated
Angelopoulos: Five-model selection bias is tiny compared to voter variability
“We don't do that right now, partially because we kind of have know from simulations that the amount of selection bias you incur with these five things is just not huge. It's not huge in comparison to the variability that you get from the, from just regular hum…”
Anastasios Angelopoulos Nov 1, 2024 ▶ 31:47 In the Arena: How LMSys changed LLM Benchmarking Forever
Nov 1, 2024 neutral
Disclosure
Angelopoulos: LMSYS considers default style control but avoids imposing opinions
“We consider that we're still actively considering it. It's just, you know, once you make that step, once you take that step, you're introducing your opinion. And I'm not, you know, why should our opinion be the one? That's kind of a community choice. We could …”
Anastasios Angelopoulos Nov 1, 2024 ▶ 18:27 In the Arena: How LMSys changed LLM Benchmarking Forever
Nov 1, 2024 neutral
Assertion Supported
Angelopoulos: The Chatbot Arena leaderboard is currently not an apples-to-apples comparison
“None of the leaderboard currently is apples to apples, because you have, like, Gemini Flash, you have, you know, all sorts of tiny models, like Llama Like, eight B and four or five B are not apples to apples.”
Anastasios Angelopoulos Nov 1, 2024 ▶ 28:03 In the Arena: How LMSys changed LLM Benchmarking Forever
Nov 1, 2024 bullish
Assertion Supported
Angelopoulos: OpenAI o1 crushed Chatbot Arena, proving the benchmark isn't saturated
“So there's this model and it crushed the benchmark. You know, it's just like really like a big gap. And what that's telling us is that it's not saturated yet. And so it's still measuring some signal that was encouraging point.”
Anastasios Angelopoulos Nov 1, 2024 ▶ 27:20 In the Arena: How LMSys changed LLM Benchmarking Forever
Nov 1, 2024 neutral
Assertion Not checkable as stated
Chiang: Chatbot Arena almost died after launch due to low engagement
“At some point, almost died. Because as you can imagine, this leaderboard depends on user, like part of, like community engagement participation. If no one comes to vote, Tomorrow then no deal.”
Wei-Lin Chiang Nov 1, 2024 ▶ 11:10 In the Arena: How LMSys changed LLM Benchmarking Forever
Nov 1, 2024 positive
Insight
Angelopoulos: Live voter data asymptotically eliminates pre-release ELO bias
“What happened is that over time, because we're getting new data, it'll get adjusted down. So if there's any bias that gets introduced at that stage in the long run, it actually doesn't matter because asymptotically, basically like in the long run, there's way …”
Anastasios Angelopoulos Nov 1, 2024 ▶ 32:16 In the Arena: How LMSys changed LLM Benchmarking Forever
Jul 31, 2025
Insight
Lambert: RLHF can never be permanently solved
“In the same way that chatbot arena can never be saturated. RLHF can never be solved.”
Nathan Lambert Jul 31, 2025 ▶ 16:42 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.