LMSYS Chatbot Arena, every mention

38 scenes · ← back to LMSYS Chatbot Arena

tap a year for its mentions
0030860152023202420252026episodesmentions
08152023202420252026episodes it came up in
0037.56152023202420252026episodesmentions per episode

every year anyone Shawn Wang 43Anastasios Angelopoulos 13Nathan Lambert 9Alessio Fanelli 2Vibhu Sapra 1Vasek Mlejnsky 1Tri Dao 1Thomas Scialom 1Pranav Reddy 1Nina Lopatina 1

Verbatim, from the transcripts: the passages where LMSYS Chatbot Arena comes up

loading…

Why AI Labs With Unlimited GPUs Still Fail — Anjney Midha, AMP Jun 18, 2026 · 2 mentions

Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay Jan 23, 2026 · 1 mention

  • ▶ 26:43 unnamed speaker Maybe an easy one to start with would be, a lot of people were focusing on maybe academic benchmarks two years ago, last year, maybe LM Arena, this year, Pokemon.

[State of Research Funding] Beyond NSF, Slingshots, Open Frontiers — Andy Konwinski, Laude Institute Dec 31, 2025 · 2 mentions

  • ▶ 6:45 Shawn Wang Uh, Arena, we already mentioned, we also just did an episode.
  • ▶ 16:29 Andy Konwinski We can actually identify projects that are more likely to become a Databricks or an Apache Spark or Array or an LM Arena sooner and, and, and with more confidence.

[State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang Dec 31, 2025 · 1 mention

  • ▶ 14:54 John Yang Either you build, like, a really compelling product, like Elmarina, that people have people use consistently, which is, I mean, really tricky in and of itself.

[State of Context Engineering] Agentic RAG, Context Rot, MCP, Subagents — Nina Lopatina, Contextual Dec 31, 2025 · 1 mention

[State of Evals] LMArena's $1.7B Vision — Anastasios Angelopoulos, LMArena Dec 31, 2025 · 37 mentions

Greg Brockman on OpenAI's Road to AGI Aug 15, 2025 · 1 mention

  • ▶ 36:45 Greg Brockman One thing I think is very interesting about these models is that we have all these arenas now, right, like LM Arena and, and others,

The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai) Jul 31, 2025 · 9 mentions

  • ▶ 12:22 Nathan Lambert It might be most of the benefit is on the, what's the right adjective to describe chatbot arena? 2 times in the scene
  • ▶ 12:48 Shawn Wang Uh, your quick, I mean, since we're there, you mentioned Sikofancy, you mentioned LM Arena. 6 times in the scene
  • ▶ 16:42 Nathan Lambert More interdisciplinary in the same way that chatbot arena can never be saturated.

⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo Jul 14, 2025 · 2 mentions

  • ▶ 1:04 Shawn Wang I think Ella Marina also had some controversies, and I think more generally people want agent evals anyway, where they are evaluating the ability of, uh, you know, models to do real tasks, and instead of like testing knowledge, they're… 2 times in the scene

⚡️Multi-Turn RL for Multi-Hour Agents — with Will Brown, Prime Intellect May 23, 2025 · 1 mention

Why Every Agent needs Open Source Cloud Sandboxes Apr 24, 2025 · 1 mention

  • ▶ 57:20 Vasek Mlejnsky A different use case, but we, for example, work with LM Arena folks from, from Berkeley that are using us to compare models in AI app generation, and we run the AI-generated app.

Gemini 2.0 Flash and Flash Thinking: the new SOTA models for the agentic era Feb 28, 2025 · 1 mention

  • ▶ 13:49 Logan Kilpatrick You have your own LMSS arena leaderboard of like people voting on the best version of the, of the newsletter on a given day would be super cool.

2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents Jan 1, 2025 · 1 mention

  • ▶ 49:50 Shawn Wang So they're trying to take on LM arena, Anastasios and crew, and they have an image arena.

Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024] Dec 24, 2024 · 1 mention

  • ▶ 17:04 Loubna Ben Allal For example, Lama 3.21 B it matches Lama two 13 B from that was the release last year on the LMSS arena, which is basically the default go to leaderboard for evaluating models using human evaluation.

The State of AI Startups in 2024 [LS Live @ NeurIPS] Dec 21, 2024 · 1 mention

In the Arena: How LMSys changed LLM Benchmarking Forever Nov 1, 2024 · 6 mentions

  • ▶ 25:51 Shawn Wang You know, you're not just running Chatbot Arena. 5 times in the scene
  • ▶ 37:45 Shawn Wang Your approach is really interesting compared to the commercial approaches where you use information from the chat arena to inform your model, which is, I mean, smart, and it's the foundation of everything you do.

[Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz Oct 13, 2024 · 1 mention

Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind Aug 28, 2024 · 1 mention

  • ▶ 48:08 unnamed speaker And this is also, by the way, the problem with Elimsis Arena, right, where, where the vast majority of prompts are single question, single answer, eval, done.

Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI Jul 23, 2024 · 3 mentions

  • ▶ 25:31 Thomas Scialom It leads quickly to, instantly to, like, state-of-the-art results for the model size, almost competing with GPT-IV on the arena leaderboard.
  • ▶ 39:01 Alessio Fanelli And I know that for example, to improve like maybe an arena score, you need different than like an MMLU score. 2 times in the scene

The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka Jul 5, 2024 · 2 mentions

  • ▶ 1:25:11 unnamed speaker One is, LMS is judge, and then two is arena, it's arena style. 2 times in the scene

State of the Art: Training 70B LLMs on 10,000 H100 clusters Jun 25, 2024 · 1 mention

How AI is Eating Finance - with Mike Conover of Brightwave Jun 11, 2024 · 1 mention

  • ▶ 1:03:03 unnamed speaker It's like, they're at the bottom of the LMC's litter board.

LLM Asia Paper Club Survey Round May 22, 2024 · 1 mention

  • ▶ 26:01 unnamed speaker Uh, we're trying to predict how people, uh, rank the, uh, chatbot responses on the arena.

The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert Jan 11, 2024 · 8 mentions

  • ▶ 52:21 unnamed speaker Um, what they're running with Chat Arena is perhaps a good store of human preference data? 2 times in the scene
  • ▶ 1:21:38 Nathan Lambert It's very hard to do if you're an engineer or a researcher, because you have your specific thing that you're zoomed in on, and it feels like a waste of time to just go play with ChatGPT or go play with ChatArena, but I really don't think… 6 times in the scene

FlashAttention-2: Making Transformers 800% faster AND exact Aug 3, 2023 · 1 mention

  • ▶ 30:29 Tri Dao Like they, they set up this kind of chat bot arena, um, to, to, to essentially benchmark different models.
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.