LMSYS Chatbot Arena, every mention
38 scenes · ← back to LMSYS Chatbot Arena
tap a year for its mentions
every year anyone Shawn Wang 43Anastasios Angelopoulos 13Nathan Lambert 9Alessio Fanelli 2Vibhu Sapra 1Vasek Mlejnsky 1Tri Dao 1Thomas Scialom 1Pranav Reddy 1Nina Lopatina 1
Verbatim, from the transcripts: the passages where LMSYS Chatbot Arena comes up
Why AI Labs With Unlimited GPUs Still Fail — Anjney Midha, AMP
- ▶ 35:17 Shawn Wang You were founding CEO of Arena? 2 times in the scene
Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay
- ▶ 26:43 unnamed speaker Maybe an easy one to start with would be, a lot of people were focusing on maybe academic benchmarks two years ago, last year, maybe LM Arena, this year, Pokemon.
[State of Research Funding] Beyond NSF, Slingshots, Open Frontiers — Andy Konwinski, Laude Institute
- ▶ 6:45 Shawn Wang Uh, Arena, we already mentioned, we also just did an episode.
- ▶ 16:29 Andy Konwinski We can actually identify projects that are more likely to become a Databricks or an Apache Spark or Array or an LM Arena sooner and, and, and with more confidence.
[State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang
[State of Context Engineering] Agentic RAG, Context Rot, MCP, Subagents — Nina Lopatina, Contextual
- ▶ 22:20 Nina Lopatina I think I'm gonna do, like, initial interviews on LM Arena.
[State of Evals] LMArena's $1.7B Vision — Anastasios Angelopoulos, LMArena
- ▶ 0:13 Shawn Wang We're here with Anastasius from ARENA. 7 times in the scene
- ▶ 1:55 Anastasios Angelopoulos We also had a great grant from Sequoia, but, um, Ansh was in particular quite, quite supportive of us and, you know, gave us some resources in order to continue building out Arena before we even 3 times in the scene
- ▶ 3:47 Shawn Wang Dude, it's an arena. 10 times in the scene
- ▶ 8:37 Shawn Wang Ok, so let's go back to Arena.
- ▶ 10:24 Anastasios Angelopoulos So, the Leaderboard Illusion's a paper that critiques Alamarina. 3 times in the scene
- ▶ 15:38 Shawn Wang I want to ask about your principles running Arena. 6 times in the scene
- ▶ 19:51 Anastasios Angelopoulos Well, so first of all, I want to give a shout out to our community manager, Greg, who is doing an awesome job managing our community, whether that's on Discord, um, or on Elemarina.
- ▶ 22:04 Anastasios Angelopoulos We need you at ARENA. 6 times in the scene
Greg Brockman on OpenAI's Road to AGI
- ▶ 36:45 Greg Brockman One thing I think is very interesting about these models is that we have all these arenas now, right, like LM Arena and, and others,
The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
- ▶ 12:22 Nathan Lambert It might be most of the benefit is on the, what's the right adjective to describe chatbot arena? 2 times in the scene
- ▶ 12:48 Shawn Wang Uh, your quick, I mean, since we're there, you mentioned Sikofancy, you mentioned LM Arena. 6 times in the scene
- ▶ 16:42 Nathan Lambert More interdisciplinary in the same way that chatbot arena can never be saturated.
⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
- ▶ 1:04 Shawn Wang I think Ella Marina also had some controversies, and I think more generally people want agent evals anyway, where they are evaluating the ability of, uh, you know, models to do real tasks, and instead of like testing knowledge, they're… 2 times in the scene
⚡️Multi-Turn RL for Multi-Hour Agents — with Will Brown, Prime Intellect
- ▶ 22:28 Shawn Wang What's your quick take on LM Arena getting a hundred million dollars?
Why Every Agent needs Open Source Cloud Sandboxes
- ▶ 57:20 Vasek Mlejnsky A different use case, but we, for example, work with LM Arena folks from, from Berkeley that are using us to compare models in AI app generation, and we run the AI-generated app.
Gemini 2.0 Flash and Flash Thinking: the new SOTA models for the agentic era
- ▶ 13:49 Logan Kilpatrick You have your own LMSS arena leaderboard of like people voting on the best version of the, of the newsletter on a given day would be super cool.
2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
- ▶ 49:50 Shawn Wang So they're trying to take on LM arena, Anastasios and crew, and they have an image arena.
Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
- ▶ 17:04 Loubna Ben Allal For example, Lama 3.21 B it matches Lama two 13 B from that was the release last year on the LMSS arena, which is basically the default go to leaderboard for evaluating models using human evaluation.
The State of AI Startups in 2024 [LS Live @ NeurIPS]
- ▶ 4:06 Pranav Reddy So this is Elm Arena.
In the Arena: How LMSys changed LLM Benchmarking Forever
- ▶ 25:51 Shawn Wang You know, you're not just running Chatbot Arena. 5 times in the scene
- ▶ 37:45 Shawn Wang Your approach is really interesting compared to the commercial approaches where you use information from the chat arena to inform your model, which is, I mean, smart, and it's the foundation of everything you do.
[Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
- ▶ 38:46 Vibhu Sapra Their ELO rankings are three X more than chatbot arena.
Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
- ▶ 48:08 unnamed speaker And this is also, by the way, the problem with Elimsis Arena, right, where, where the vast majority of prompts are single question, single answer, eval, done.
Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
- ▶ 25:31 Thomas Scialom It leads quickly to, instantly to, like, state-of-the-art results for the model size, almost competing with GPT-IV on the arena leaderboard.
- ▶ 39:01 Alessio Fanelli And I know that for example, to improve like maybe an arena score, you need different than like an MMLU score. 2 times in the scene
The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
- ▶ 1:25:11 unnamed speaker One is, LMS is judge, and then two is arena, it's arena style. 2 times in the scene
State of the Art: Training 70B LLMs on 10,000 H100 clusters
- ▶ 1:29:37 Jonathan Frankle Um, cool stuff kind of popping up sometimes on chatbot arena and, you know, keep your eyes on.
How AI is Eating Finance - with Mike Conover of Brightwave
- ▶ 1:03:03 unnamed speaker It's like, they're at the bottom of the LMC's litter board.
LLM Asia Paper Club Survey Round
- ▶ 26:01 unnamed speaker Uh, we're trying to predict how people, uh, rank the, uh, chatbot responses on the arena.
The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
- ▶ 52:21 unnamed speaker Um, what they're running with Chat Arena is perhaps a good store of human preference data? 2 times in the scene
- ▶ 1:21:38 Nathan Lambert It's very hard to do if you're an engineer or a researcher, because you have your specific thing that you're zoomed in on, and it feels like a waste of time to just go play with ChatGPT or go play with ChatArena, but I really don't think… 6 times in the scene