LMSYS Chatbot Arena

part of LMSYS

4 statements across 4 episodes · 3 bullish · 0 bearish · 4 people on the record · first statement Jan 11, 2024 by Nathan Lambert · said 87 times in 25 episodes since 2023 · across every show →

Mentions by year

brought up most by Shawn Wang (43), Anastasios Angelopoulos (13), Nathan Lambert (9), Alessio Fanelli (2), Vibhu Sapra (1), Vasek Mlejnsky (1), Tri Dao (1), Thomas Scialom (1)

tap a year for its mentions
0030860152023202420252026episodesmentions
08152023202420252026episodes it came up in
0037.56152023202420252026episodesmentions per episode
2026 3 mentions in 2 episodes 2 per episode
2025 57 mentions in 11 episodes 5 per episode
2024 26 mentions in 11 episodes 2 per episode
2023 1 mention in 1 episode

every mention, scene by scene, with the transcript →

Everything said about LMSYS Chatbot Arena, oldest first

Jan 11, 2024 positive
Assertion Supported
Lambert: GPT-4 Turbo Showed a Noticeable Jump on LMSYS Chatbot Arena
“GPT-IV Turbo is also notably ahead of the other GPT-IVs, which it kind of showed up immediately once they added it to the leaderboard, or to the arena, and I was like, all the GPT-IV memes aside, it seems like this is effectively a bump in the model.”
Nathan Lambert Jan 11, 2024 ▶ 1:25:01 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 1, 2025 positive
Assertion Supported
Swyx: Iso-ELO LLM token costs dropped nearly 1,000x in 2024
“For the same amount of ELO, what you used to pay at the start of 24 you know, let's say, you know, 50, 40 to 50 dollars Per million tokens now is available approximately at, with Amazon Nova approximately at, I don't know, 0.075 dollars per token. So like seve…”
Shawn Wang Jan 1, 2025 ▶ 1:20:51 2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
Jan 19, 2025 bullish
Opinion
Zhang: DeepSeek-V3 is currently the leading open-source LLM
“Yeah, because DeepSeq VIII, I think, is currently considered the leading open source LLMs based on the benchmark results and the chat area results.”
Yining Zhang Jan 19, 2025 ▶ 1:22 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Dec 31, 2025 neutral
Insight
Yang: Evaluating Human-AI Code Interaction Requires Compelling Products or Simulators
“From an academic standpoint, it feels like there's two difficult approaches to resolving that. Either you build, like, a really compelling product, like Elmarina, that people have people use consistently, which is, I mean, really tricky in and of itself. Or yo…”
John Yang Dec 31, 2025 ▶ 14:48 [State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.