May 29, 2025 · 1h 45m · a16z

Beyond Leaderboards: LMArena’s Mission to Make AI Reliable

Anastasios Angelopoulos · 28m spoken Ion Stoica · 27m spoken Wei-Lin Chiang · 18m spoken Anjney Midha · 17m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the a16z podcast, host Anjney Midha converses with LMArena co-founders Anastasios Angelopoulos, Wei-Lin Chiang, and Ion Stoica about transforming AI evaluation from static benchmarks into real-time, crowdsourced human preference testing. They explore LMArena's academic origins at UC Berkeley, technical innovations like Style Control and dynamic routing, and its evolution into an open, mission-critical testing harness for frontier AI models.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The host as informed peer 2.6 Guest teaching 3.9 Guest disagreement 0.7 The host pushing back 1.2
05100:0020:0040:001:00:001:20:001:40:000:00–3:52 · The host as informed peer 3/10 Title Sequence: Beyond Leaderboards and AI Reliability The host asks informed questions about scaling evaluation to mission-critical industries. Anastasios politely reframes the premise by explaining that hard science questions are mostly subjective rather than simple factual lookups.3:52–6:04 · The host as informed peer 2/10 Pre-Release Testing and Model Release Pipelines The host asks whether LMArena assists large labs over open-source labs. Anastasios explains their release pipeline and pre-release testing for all developers.6:04–14:22 · The host as informed peer 4/10 Crowdsourcing Expertise and Style Control The host intentionally channels expert criticism, arguing that lay users prefer slop while experts should set standards. The guests push back strongly, showing that top experts lack time to label and introducing style control to adjust for user biases statistically.14:22–18:51 · The host as informed peer 2/10 Expanding Beyond Chatbots: Web Dev Arena The host asks why Chatbot Arena was insufficient for coding, leading to Web Dev Arena. Guests explain the need for real-time sandbox execution and browser feedback.18:51–22:45 · The host as informed peer 3/10 Overcoming Benchmark Contamination and Overfitting The host and guests discuss benchmark contamination and overfitting. The guests explain how continuous data collection creates an operational defense against memorization.22:45–29:50 · The host as informed peer 3/10 Personalization and Hard Signal in Real-World Testing The host asks why Web Dev Arena serves as a general capability proxy, implying chat might be easier. Anastasios strongly disagrees with the premise that chat is easy, calling it naive and highlighting the deep subjective complexity of conversation.29:50–39:20 · The host as informed peer 2/10 The Roots of LMArena: From Vicuna to Elo The host prompts the guests to share the origin story of LMArena from the Vicuna project. Guests detail how they moved from student pizza labeling to LLM-as-judge and ultimately to Elo and Bradley-Terry statistical models.39:20–41:31 · The host as informed peer 2/10 University Research and Academic Neutrality The host asks if Arena could only originate at an interdisciplinary university lab. Guests emphasize academic neutrality and small, nimble cross-disciplinary teams as key advantages.41:31–48:30 · The host as informed peer 3/10 Proving the Value of Academic AI Research The host recalls the early 2023 narrative declaring the death of frontier AI research in academia. Ion recounts how small university teams proved skeptics wrong through systems like Vicuna and VLLM.48:30–1:00:02 · The host as informed peer 3/10 Scaling LMArena and Commercializing Research The host highlights platform growth statistics. Guests detail the transition to a commercial entity to scale infrastructure and explain Prompt to Leaderboard methodology.1:00:02–1:12:16 · The host as informed peer 3/10 Benchmarks vs. Evaluation in the RL Era The host asks about the fundamental difference between static benchmarks and evaluation. Anastasios educates the room by framing static benchmarks as supervised learning and human preference arena testing as reinforcement learning from the open world.1:12:16–1:18:01 · The host as informed peer 2/10 Challenges in Measuring AI Reliability and Granularity The host asks what makes granular evaluation technically difficult. Anastasios explains matrix sparsity across infinite potential queries and finite user votes.1:18:01–1:28:11 · The host as informed peer 3/10 Expanding Beyond Binary Rankings as Models Evolve The host probes how Arena handles diverging product capabilities like built-in memory. Guests explain plans for app integrations, SDKs, and data-driven debugging (D3) using implicit usage signals.1:28:11–1:31:32 · The host as informed peer 3/10 Prompt to Leaderboard and Cost-Constrained Routing The host asks about the performance of the open-source Prompt to Leaderboard repo. Anastasios details how routing prompts via estimated Bradley-Terry scores delivers a 2x cost-performance improvement over single models.1:31:32–1:34:34 · The host as informed peer 2/10 The LMArena Roadmap: Personalization and User Leaderboards The host asks for upcoming roadmap features. Guests explain personal leaderboards and user rank leaderboards that align voter incentives to reduce spam and noise.1:34:34–1:37:39 · The host as informed peer 2/10 The Importance of Open Source and Openness The host asks how the company maintains its core values post-commercialization. Guests emphasize that open code, open data, and scientific transparency are essential for building trust and attracting talent.1:37:39–1:39:40 · The host as informed peer 3/10 Balancing Open Evaluation and Mission-Critical Security The host addresses the tension between open public testing and closed mission-critical security demands. Anastasios notes that while national security is outside his domain, private deployments accommodate sensitive environments.1:39:40–1:43:16 · The host as informed peer 3/10 Red Team Arena and Security Evaluation The host asks about the operational mechanics of Red Team Arena. Wei-Lin explains how crowdsourced jailbreaking tests safety guidelines in realistic application environments.1:43:16–1:44:55 · The host as informed peer 2/10 Adapting to AI Agents and Future Evolutions The host asks if agentic long-horizon tasks will require fundamental architectural changes to Arena. Anastasios asserts that while UIs will change, real-world organic testing with human feedback remains the immutable core.0:00–3:52 · Guest teaching 5/10 Title Sequence: Beyond Leaderboards and AI Reliability The host asks informed questions about scaling evaluation to mission-critical industries. Anastasios politely reframes the premise by explaining that hard science questions are mostly subjective rather than simple factual lookups.3:52–6:04 · Guest teaching 3/10 Pre-Release Testing and Model Release Pipelines The host asks whether LMArena assists large labs over open-source labs. Anastasios explains their release pipeline and pre-release testing for all developers.6:04–14:22 · Guest teaching 6/10 Crowdsourcing Expertise and Style Control The host intentionally channels expert criticism, arguing that lay users prefer slop while experts should set standards. The guests push back strongly, showing that top experts lack time to label and introducing style control to adjust for user biases statistically.14:22–18:51 · Guest teaching 3/10 Expanding Beyond Chatbots: Web Dev Arena The host asks why Chatbot Arena was insufficient for coding, leading to Web Dev Arena. Guests explain the need for real-time sandbox execution and browser feedback.18:51–22:45 · Guest teaching 4/10 Overcoming Benchmark Contamination and Overfitting The host and guests discuss benchmark contamination and overfitting. The guests explain how continuous data collection creates an operational defense against memorization.22:45–29:50 · Guest teaching 4/10 Personalization and Hard Signal in Real-World Testing The host asks why Web Dev Arena serves as a general capability proxy, implying chat might be easier. Anastasios strongly disagrees with the premise that chat is easy, calling it naive and highlighting the deep subjective complexity of conversation.29:50–39:20 · Guest teaching 4/10 The Roots of LMArena: From Vicuna to Elo The host prompts the guests to share the origin story of LMArena from the Vicuna project. Guests detail how they moved from student pizza labeling to LLM-as-judge and ultimately to Elo and Bradley-Terry statistical models.39:20–41:31 · Guest teaching 3/10 University Research and Academic Neutrality The host asks if Arena could only originate at an interdisciplinary university lab. Guests emphasize academic neutrality and small, nimble cross-disciplinary teams as key advantages.41:31–48:30 · Guest teaching 4/10 Proving the Value of Academic AI Research The host recalls the early 2023 narrative declaring the death of frontier AI research in academia. Ion recounts how small university teams proved skeptics wrong through systems like Vicuna and VLLM.48:30–1:00:02 · Guest teaching 4/10 Scaling LMArena and Commercializing Research The host highlights platform growth statistics. Guests detail the transition to a commercial entity to scale infrastructure and explain Prompt to Leaderboard methodology.1:00:02–1:12:16 · Guest teaching 5/10 Benchmarks vs. Evaluation in the RL Era The host asks about the fundamental difference between static benchmarks and evaluation. Anastasios educates the room by framing static benchmarks as supervised learning and human preference arena testing as reinforcement learning from the open world.1:12:16–1:18:01 · Guest teaching 4/10 Challenges in Measuring AI Reliability and Granularity The host asks what makes granular evaluation technically difficult. Anastasios explains matrix sparsity across infinite potential queries and finite user votes.1:18:01–1:28:11 · Guest teaching 4/10 Expanding Beyond Binary Rankings as Models Evolve The host probes how Arena handles diverging product capabilities like built-in memory. Guests explain plans for app integrations, SDKs, and data-driven debugging (D3) using implicit usage signals.1:28:11–1:31:32 · Guest teaching 5/10 Prompt to Leaderboard and Cost-Constrained Routing The host asks about the performance of the open-source Prompt to Leaderboard repo. Anastasios details how routing prompts via estimated Bradley-Terry scores delivers a 2x cost-performance improvement over single models.1:31:32–1:34:34 · Guest teaching 3/10 The LMArena Roadmap: Personalization and User Leaderboards The host asks for upcoming roadmap features. Guests explain personal leaderboards and user rank leaderboards that align voter incentives to reduce spam and noise.1:34:34–1:37:39 · Guest teaching 3/10 The Importance of Open Source and Openness The host asks how the company maintains its core values post-commercialization. Guests emphasize that open code, open data, and scientific transparency are essential for building trust and attracting talent.1:37:39–1:39:40 · Guest teaching 3/10 Balancing Open Evaluation and Mission-Critical Security The host addresses the tension between open public testing and closed mission-critical security demands. Anastasios notes that while national security is outside his domain, private deployments accommodate sensitive environments.1:39:40–1:43:16 · Guest teaching 4/10 Red Team Arena and Security Evaluation The host asks about the operational mechanics of Red Team Arena. Wei-Lin explains how crowdsourced jailbreaking tests safety guidelines in realistic application environments.1:43:16–1:44:55 · Guest teaching 3/10 Adapting to AI Agents and Future Evolutions The host asks if agentic long-horizon tasks will require fundamental architectural changes to Arena. Anastasios asserts that while UIs will change, real-world organic testing with human feedback remains the immutable core.0:00–3:52 · Guest disagreement 2/10 Title Sequence: Beyond Leaderboards and AI Reliability The host asks informed questions about scaling evaluation to mission-critical industries. Anastasios politely reframes the premise by explaining that hard science questions are mostly subjective rather than simple factual lookups.3:52–6:04 · Guest disagreement 1/10 Pre-Release Testing and Model Release Pipelines The host asks whether LMArena assists large labs over open-source labs. Anastasios explains their release pipeline and pre-release testing for all developers.6:04–14:22 · Guest disagreement 3/10 Crowdsourcing Expertise and Style Control The host intentionally channels expert criticism, arguing that lay users prefer slop while experts should set standards. The guests push back strongly, showing that top experts lack time to label and introducing style control to adjust for user biases statistically.14:22–18:51 · Guest disagreement 0/10 Expanding Beyond Chatbots: Web Dev Arena The host asks why Chatbot Arena was insufficient for coding, leading to Web Dev Arena. Guests explain the need for real-time sandbox execution and browser feedback.18:51–22:45 · Guest disagreement 1/10 Overcoming Benchmark Contamination and Overfitting The host and guests discuss benchmark contamination and overfitting. The guests explain how continuous data collection creates an operational defense against memorization.22:45–29:50 · Guest disagreement 3/10 Personalization and Hard Signal in Real-World Testing The host asks why Web Dev Arena serves as a general capability proxy, implying chat might be easier. Anastasios strongly disagrees with the premise that chat is easy, calling it naive and highlighting the deep subjective complexity of conversation.29:50–39:20 · Guest disagreement 0/10 The Roots of LMArena: From Vicuna to Elo The host prompts the guests to share the origin story of LMArena from the Vicuna project. Guests detail how they moved from student pizza labeling to LLM-as-judge and ultimately to Elo and Bradley-Terry statistical models.39:20–41:31 · Guest disagreement 0/10 University Research and Academic Neutrality The host asks if Arena could only originate at an interdisciplinary university lab. Guests emphasize academic neutrality and small, nimble cross-disciplinary teams as key advantages.41:31–48:30 · Guest disagreement 1/10 Proving the Value of Academic AI Research The host recalls the early 2023 narrative declaring the death of frontier AI research in academia. Ion recounts how small university teams proved skeptics wrong through systems like Vicuna and VLLM.48:30–1:00:02 · Guest disagreement 0/10 Scaling LMArena and Commercializing Research The host highlights platform growth statistics. Guests detail the transition to a commercial entity to scale infrastructure and explain Prompt to Leaderboard methodology.1:00:02–1:12:16 · Guest disagreement 2/10 Benchmarks vs. Evaluation in the RL Era The host asks about the fundamental difference between static benchmarks and evaluation. Anastasios educates the room by framing static benchmarks as supervised learning and human preference arena testing as reinforcement learning from the open world.1:12:16–1:18:01 · Guest disagreement 0/10 Challenges in Measuring AI Reliability and Granularity The host asks what makes granular evaluation technically difficult. Anastasios explains matrix sparsity across infinite potential queries and finite user votes.1:18:01–1:28:11 · Guest disagreement 0/10 Expanding Beyond Binary Rankings as Models Evolve The host probes how Arena handles diverging product capabilities like built-in memory. Guests explain plans for app integrations, SDKs, and data-driven debugging (D3) using implicit usage signals.1:28:11–1:31:32 · Guest disagreement 0/10 Prompt to Leaderboard and Cost-Constrained Routing The host asks about the performance of the open-source Prompt to Leaderboard repo. Anastasios details how routing prompts via estimated Bradley-Terry scores delivers a 2x cost-performance improvement over single models.1:31:32–1:34:34 · Guest disagreement 0/10 The LMArena Roadmap: Personalization and User Leaderboards The host asks for upcoming roadmap features. Guests explain personal leaderboards and user rank leaderboards that align voter incentives to reduce spam and noise.1:34:34–1:37:39 · Guest disagreement 0/10 The Importance of Open Source and Openness The host asks how the company maintains its core values post-commercialization. Guests emphasize that open code, open data, and scientific transparency are essential for building trust and attracting talent.1:37:39–1:39:40 · Guest disagreement 1/10 Balancing Open Evaluation and Mission-Critical Security The host addresses the tension between open public testing and closed mission-critical security demands. Anastasios notes that while national security is outside his domain, private deployments accommodate sensitive environments.1:39:40–1:43:16 · Guest disagreement 0/10 Red Team Arena and Security Evaluation The host asks about the operational mechanics of Red Team Arena. Wei-Lin explains how crowdsourced jailbreaking tests safety guidelines in realistic application environments.1:43:16–1:44:55 · Guest disagreement 0/10 Adapting to AI Agents and Future Evolutions The host asks if agentic long-horizon tasks will require fundamental architectural changes to Arena. Anastasios asserts that while UIs will change, real-world organic testing with human feedback remains the immutable core.0:00–3:52 · The host pushing back 1/10 Title Sequence: Beyond Leaderboards and AI Reliability The host asks informed questions about scaling evaluation to mission-critical industries. Anastasios politely reframes the premise by explaining that hard science questions are mostly subjective rather than simple factual lookups.3:52–6:04 · The host pushing back 1/10 Pre-Release Testing and Model Release Pipelines The host asks whether LMArena assists large labs over open-source labs. Anastasios explains their release pipeline and pre-release testing for all developers.6:04–14:22 · The host pushing back 6/10 Crowdsourcing Expertise and Style Control The host intentionally channels expert criticism, arguing that lay users prefer slop while experts should set standards. The guests push back strongly, showing that top experts lack time to label and introducing style control to adjust for user biases statistically.14:22–18:51 · The host pushing back 1/10 Expanding Beyond Chatbots: Web Dev Arena The host asks why Chatbot Arena was insufficient for coding, leading to Web Dev Arena. Guests explain the need for real-time sandbox execution and browser feedback.18:51–22:45 · The host pushing back 1/10 Overcoming Benchmark Contamination and Overfitting The host and guests discuss benchmark contamination and overfitting. The guests explain how continuous data collection creates an operational defense against memorization.22:45–29:50 · The host pushing back 2/10 Personalization and Hard Signal in Real-World Testing The host asks why Web Dev Arena serves as a general capability proxy, implying chat might be easier. Anastasios strongly disagrees with the premise that chat is easy, calling it naive and highlighting the deep subjective complexity of conversation.29:50–39:20 · The host pushing back 0/10 The Roots of LMArena: From Vicuna to Elo The host prompts the guests to share the origin story of LMArena from the Vicuna project. Guests detail how they moved from student pizza labeling to LLM-as-judge and ultimately to Elo and Bradley-Terry statistical models.39:20–41:31 · The host pushing back 1/10 University Research and Academic Neutrality The host asks if Arena could only originate at an interdisciplinary university lab. Guests emphasize academic neutrality and small, nimble cross-disciplinary teams as key advantages.41:31–48:30 · The host pushing back 1/10 Proving the Value of Academic AI Research The host recalls the early 2023 narrative declaring the death of frontier AI research in academia. Ion recounts how small university teams proved skeptics wrong through systems like Vicuna and VLLM.48:30–1:00:02 · The host pushing back 1/10 Scaling LMArena and Commercializing Research The host highlights platform growth statistics. Guests detail the transition to a commercial entity to scale infrastructure and explain Prompt to Leaderboard methodology.1:00:02–1:12:16 · The host pushing back 2/10 Benchmarks vs. Evaluation in the RL Era The host asks about the fundamental difference between static benchmarks and evaluation. Anastasios educates the room by framing static benchmarks as supervised learning and human preference arena testing as reinforcement learning from the open world.1:12:16–1:18:01 · The host pushing back 1/10 Challenges in Measuring AI Reliability and Granularity The host asks what makes granular evaluation technically difficult. Anastasios explains matrix sparsity across infinite potential queries and finite user votes.1:18:01–1:28:11 · The host pushing back 2/10 Expanding Beyond Binary Rankings as Models Evolve The host probes how Arena handles diverging product capabilities like built-in memory. Guests explain plans for app integrations, SDKs, and data-driven debugging (D3) using implicit usage signals.1:28:11–1:31:32 · The host pushing back 0/10 Prompt to Leaderboard and Cost-Constrained Routing The host asks about the performance of the open-source Prompt to Leaderboard repo. Anastasios details how routing prompts via estimated Bradley-Terry scores delivers a 2x cost-performance improvement over single models.1:31:32–1:34:34 · The host pushing back 0/10 The LMArena Roadmap: Personalization and User Leaderboards The host asks for upcoming roadmap features. Guests explain personal leaderboards and user rank leaderboards that align voter incentives to reduce spam and noise.1:34:34–1:37:39 · The host pushing back 0/10 The Importance of Open Source and Openness The host asks how the company maintains its core values post-commercialization. Guests emphasize that open code, open data, and scientific transparency are essential for building trust and attracting talent.1:37:39–1:39:40 · The host pushing back 2/10 Balancing Open Evaluation and Mission-Critical Security The host addresses the tension between open public testing and closed mission-critical security demands. Anastasios notes that while national security is outside his domain, private deployments accommodate sensitive environments.1:39:40–1:43:16 · The host pushing back 1/10 Red Team Arena and Security Evaluation The host asks about the operational mechanics of Red Team Arena. Wei-Lin explains how crowdsourced jailbreaking tests safety guidelines in realistic application environments.1:43:16–1:44:55 · The host pushing back 0/10 Adapting to AI Agents and Future Evolutions The host asks if agentic long-horizon tasks will require fundamental architectural changes to Arena. Anastasios asserts that while UIs will change, real-world organic testing with human feedback remains the immutable core.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 0% · guest 100%0:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%33:00 · the host 0% · guest 100%33:00 · the host 0% · guest 100%36:00 · the host 0% · guest 100%36:00 · the host 0% · guest 100%39:00 · the host 0% · guest 100%39:00 · the host 0% · guest 100%42:00 · the host 0% · guest 100%42:00 · the host 0% · guest 100%45:00 · the host 0% · guest 100%45:00 · the host 0% · guest 100%48:00 · the host 0% · guest 100%48:00 · the host 0% · guest 100%51:00 · the host 0% · guest 100%51:00 · the host 0% · guest 100%54:00 · the host 0% · guest 100%54:00 · the host 0% · guest 100%57:00 · the host 0% · guest 100%57:00 · the host 0% · guest 100%1:00:00 · the host 0% · guest 100%1:00:00 · the host 0% · guest 100%1:03:00 · the host 0% · guest 100%1:03:00 · the host 0% · guest 100%1:06:00 · the host 0% · guest 100%1:06:00 · the host 0% · guest 100%1:09:00 · the host 0% · guest 100%1:09:00 · the host 0% · guest 100%1:12:00 · the host 0% · guest 100%1:12:00 · the host 0% · guest 100%1:15:00 · the host 0% · guest 100%1:15:00 · the host 0% · guest 100%1:18:00 · the host 0% · guest 100%1:18:00 · the host 0% · guest 100%1:21:00 · the host 0% · guest 100%1:21:00 · the host 0% · guest 100%1:24:00 · the host 0% · guest 100%1:24:00 · the host 0% · guest 100%1:27:00 · the host 0% · guest 100%1:27:00 · the host 0% · guest 100%1:30:00 · the host 0% · guest 100%1:30:00 · the host 0% · guest 100%1:33:00 · the host 0% · guest 100%1:33:00 · the host 0% · guest 100%1:36:00 · the host 0% · guest 100%1:36:00 · the host 0% · guest 100%1:39:00 · the host 0% · guest 100%1:39:00 · the host 0% · guest 100%1:42:00 · the host 0% · guest 100%1:42:00 · the host 0% · guest 100%1:45:00 · the host 0% · guest 0%1:45:00 · the host 0% · guest 0%
Sharpest disagreement ▶ 24:56 Rejection of the assumption that chat evaluation is easy

Anastasios explicitly rejects the host's premise, calling the assumption that chat is easy or naive and forcefully explaining the deep subjective complexity involved.

Hardest push from the host ▶ 8:18 Challenging crowdsourcing with 'lay people prefer slop'

The host explicitly announces a pushback, challenging the crowdsourcing model by arguing that educated experts should set standards while lay users prefer slop.

Biggest teaching moment ▶ 1:00:42 Reframing static evaluation as Supervised vs RL

Anastasios provides a powerful conceptual reframe, showing that static benchmarks function like limited supervised learning while arena evaluations function like reinforcement learning from the open world.

The host holds their own ▶ 1:28:20 Drilling into Prompt to Leaderboard routing mechanics

The host demonstrates deep familiarity with the open-source repository and technical details, pressing the guest to explain how the router achieves superior statistical performance.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
Title Sequence: Beyond Leaderboards and AI Reliability 3521 The host asks informed questions about scaling evaluation to mission-critical industries. Anastasios politely reframes the premise by explaining that hard science questions are mostly subjective rather than simple factual lookups.
Pre-Release Testing and Model Release Pipelines 2311 The host asks whether LMArena assists large labs over open-source labs. Anastasios explains their release pipeline and pre-release testing for all developers.
Crowdsourcing Expertise and Style Control 4636 The host intentionally channels expert criticism, arguing that lay users prefer slop while experts should set standards. The guests push back strongly, showing that top experts lack time to label and introducing style control to adjust for user biases statistically.
Expanding Beyond Chatbots: Web Dev Arena 2301 The host asks why Chatbot Arena was insufficient for coding, leading to Web Dev Arena. Guests explain the need for real-time sandbox execution and browser feedback.
Overcoming Benchmark Contamination and Overfitting 3411 The host and guests discuss benchmark contamination and overfitting. The guests explain how continuous data collection creates an operational defense against memorization.
Personalization and Hard Signal in Real-World Testing 3432 The host asks why Web Dev Arena serves as a general capability proxy, implying chat might be easier. Anastasios strongly disagrees with the premise that chat is easy, calling it naive and highlighting the deep subjective complexity of conversation.
The Roots of LMArena: From Vicuna to Elo 2400 The host prompts the guests to share the origin story of LMArena from the Vicuna project. Guests detail how they moved from student pizza labeling to LLM-as-judge and ultimately to Elo and Bradley-Terry statistical models.
University Research and Academic Neutrality 2301 The host asks if Arena could only originate at an interdisciplinary university lab. Guests emphasize academic neutrality and small, nimble cross-disciplinary teams as key advantages.
Proving the Value of Academic AI Research 3411 The host recalls the early 2023 narrative declaring the death of frontier AI research in academia. Ion recounts how small university teams proved skeptics wrong through systems like Vicuna and VLLM.
Scaling LMArena and Commercializing Research 3401 The host highlights platform growth statistics. Guests detail the transition to a commercial entity to scale infrastructure and explain Prompt to Leaderboard methodology.
Benchmarks vs. Evaluation in the RL Era 3522 The host asks about the fundamental difference between static benchmarks and evaluation. Anastasios educates the room by framing static benchmarks as supervised learning and human preference arena testing as reinforcement learning from the open world.
Challenges in Measuring AI Reliability and Granularity 2401 The host asks what makes granular evaluation technically difficult. Anastasios explains matrix sparsity across infinite potential queries and finite user votes.
Expanding Beyond Binary Rankings as Models Evolve 3402 The host probes how Arena handles diverging product capabilities like built-in memory. Guests explain plans for app integrations, SDKs, and data-driven debugging (D3) using implicit usage signals.
Prompt to Leaderboard and Cost-Constrained Routing 3500 The host asks about the performance of the open-source Prompt to Leaderboard repo. Anastasios details how routing prompts via estimated Bradley-Terry scores delivers a 2x cost-performance improvement over single models.
The LMArena Roadmap: Personalization and User Leaderboards 2300 The host asks for upcoming roadmap features. Guests explain personal leaderboards and user rank leaderboards that align voter incentives to reduce spam and noise.
The Importance of Open Source and Openness 2300 The host asks how the company maintains its core values post-commercialization. Guests emphasize that open code, open data, and scientific transparency are essential for building trust and attracting talent.
Balancing Open Evaluation and Mission-Critical Security 3312 The host addresses the tension between open public testing and closed mission-critical security demands. Anastasios notes that while national security is outside his domain, private deployments accommodate sensitive environments.
Red Team Arena and Security Evaluation 3401 The host asks about the operational mechanics of Red Team Arena. Wei-Lin explains how crowdsourced jailbreaking tests safety guidelines in realistic application environments.
Adapting to AI Agents and Future Evolutions 2300 The host asks if agentic long-horizon tasks will require fundamental architectural changes to Arena. Anastasios asserts that while UIs will change, real-world organic testing with human feedback remains the immutable core.

Statements from this episode (31)

Opinion
Midha: LMArena serves as humanity's real-time exam for AI models
“It sounds like what Irina is, is humanity's real-time exam.”
Anjney Midha May 29, 2025 ▶ 0:08
Prediction Not checkable as stated
Midha: Real-time testing will replace static AI benchmarks like MMLU
“While benchmarks like MMLU and the idea of these static exams were useful three years ago, the future is about real-time evaluation, real-time systems, real-time testing in the wild.”
Anjney Midha May 29, 2025 ▶ 0:43
Insight
Angelopoulos: Most mission-critical AI queries are subjective, not factual lookups
“In reality, even in such industries, the majority of questions that people ask are subjective. Okay. So the mythology that in hard sciences or in mission critical industries, people just have like cut and dried questions and they just need like a retrieval and…”
Anastasios Angelopoulos May 29, 2025 ▶ 2:50
Disclosure
Angelopoulos: LMArena conducts pre-release model testing for AI developers
“One of the things that we help everybody to do is pre-release testing of their models. Okay. So it's not just that, you know, we work together to evaluate the models are released, but we also try to be their release partners and say, Hey, can we help you guys …”
Anastasios Angelopoulos May 29, 2025 ▶ 4:38
Assertion Not checkable as stated
Stoica: Top technical experts lack time to perform AI evaluation labeling
“You're getting some people from that area who are willing to do the labeling, but the best people are not willing. Fundamentally, they don't have time.”
Ion Stoica May 29, 2025 ▶ 9:51
Assertion Open · timeframe May 2026
Angelopoulos: Human evaluators prefer longer AI responses given equal content
“It's true that people vote for longer responses, you know, preferentially over shorter responses, even given the same contents or well-known human bias.”
Anastasios Angelopoulos May 29, 2025 ▶ 12:31
Assertion Supported
Angelopoulos: LMArena makes style control the default AI evaluation method
“That's why we're making style control default.”
Anastasios Angelopoulos May 29, 2025 ▶ 12:55
Assertion Not checkable as stated
Angelopoulos: LMArena measures user preference, not AGI progress
“We don't claim to be an AGI benchmark. We are faithfully representing the preferences of our community.”
Anastasios Angelopoulos May 29, 2025 ▶ 18:08
Assertion Not checkable as stated
Stoica: Over 70% of daily LMArena prompts are completely unique
“And basically measures out how many more, you know, fresh prompts you have in one day compared to what you've seen in the past three months, right? And by a similarity score of something like 70, 75%, you have over 70 of these prompts are fresh.”
Ion Stoica May 29, 2025 ▶ 20:13
Assertion Not checkable as stated
Angelopoulos: Chatbot Arena is immune to model overfitting by design
“Static benchmarks overfit. Why? It's because as Jan said earlier, you're giving the student the same test over and over. You have a model, you test it, you know, you look at whether or not it's improved on a static data set. Then you find another model, you te…”
Anastasios Angelopoulos May 29, 2025 ▶ 21:00
Opinion
Angelopoulos: Believing chatbot leaderboards are easily gameable is naive
“I have to say, I also just like completely disagree with the foundation of the question. The like implicit assumption is that like chat is easy or that it's even easier than web dev. That's completely false. It's a completely naive perspective that people have…”
Anastasios Angelopoulos May 29, 2025 ▶ 24:57
Disclosure
Angelopoulos: LMArena is building personalized AI leaderboard tools
“Yeah, and we should be giving you the tools to do that, and we're currently building them.”
Anastasios Angelopoulos May 29, 2025 ▶ 25:56
Prediction Not checkable as stated
Angelopoulos: Future AI evaluation will shift to personalized user leaderboards
“Absolutely. Absolutely. It should be personalized just for you. You should understand which models are best for you.”
Anastasios Angelopoulos May 29, 2025 ▶ 26:06
Assertion Supported
Angelopoulos: Bradley-Terry models converge for AI evaluation, unlike Elo scores
“Okay, let's move from Elo to Bradley Terry because we're actually performing an estimate here instead of just like You know, and the ELO score moves over time. It doesn't converge, but Rally Terry models converge and how do we then construct confidence interva…”
Anastasios Angelopoulos May 29, 2025 ▶ 38:49
Opinion
Angelopoulos: Industry AI evaluation platforms face skepticism over bias
“The fact that we come from Berkeley and from a university really speaks to our scientific approach in neutrality. I think if it came from an industrial lab, people would always have questions about, oh, well, these people are they also training a model and wha…”
Anastasios Angelopoulos May 29, 2025 ▶ 39:36
Assertion Supported
Midha: LMArena hosts over 280 AI models, up from 12 initially
“In total, that first year there were about 12 models or so, and today it's over 280 or something on the platform.”
Anjney Midha May 29, 2025 ▶ 49:05
Insight
Angelopoulos: AI evaluation performance follows a data scaling law
“Because language models are sort of the intermediary that gets you to this evaluation, there's also a scaling law that comes along with it. Which is to say that the more data you get, the bigger you build the platform, the better you can make your evaluations,…”
Anastasios Angelopoulos May 29, 2025 ▶ 56:18
Insight
Angelopoulos: Reinforcement learning allows AI models to surpass human teachers
“And supervised learning, you can only do as well as the best human that you have. Because what's happening is that you're learning from the teacher. In reinforcement learning, you're learning from the world. You're able to learn things better than the best hum…”
Anastasios Angelopoulos May 29, 2025 ▶ 1:00:48
Assertion Open · timeframe May 2026
Stoica: Users grading answers to their own questions is the gold standard
“This is now from information retrieval field for decades, and it's called gold standard, when people evaluate the answer to their own questions. When an expert evaluates someone else questions, and the answer is called the silver, if I remember correctly.”
Ion Stoica May 29, 2025 ▶ 1:08:26
Assertion Not checkable as stated
Angelopoulos: Chatbot Arena has 1M+ monthly users and 150M+ conversations
“A lot of people don't know this, but ShopBot Arena is Used by like a million plus monthly users. We get like, you know, tens of thousands of votes on a daily basis. We have like over like, you know, a hundred fifty million conversations that have been had on t…”
Anastasios Angelopoulos May 29, 2025 ▶ 1:13:19
Prediction Didn’t hold up
Angelopoulos: LMArena will launch Data-Driven Debugging within months
“So we're building a project now that we call data-driven debugging D three. It's, you know, it's a little farther out. It'll come in a couple months.”
Anastasios Angelopoulos May 29, 2025 ▶ 1:26:14
Insight
Angelopoulos: AI leaderboards can utilize any form of interaction feedback
“Pairwise comparison feedback is not the only kind of feedback that we can use to construct leaderboards. We can construct leaderboards with any form of feedback.”
Anastasios Angelopoulos May 29, 2025 ▶ 1:26:25
Assertion Supported
Angelopoulos: LMArena router model outperforms all constituent models on Chatbot Arena
“When you train a prompt to leaderboard model, which is like, let's say a seven billion parameter model, and then you use it to route on just questions on the arena and everybody's questions, that model does better than any of the constituent models that were u…”
Anastasios Angelopoulos May 29, 2025 ▶ 1:29:10
Assertion Supported
Angelopoulos: LMArena prompt router yields double the performance per dollar
“Now, if you trace the performance, the best performance that, you know, any individual model can give you as part of the router as a function of cost. That's like two X worse than the router. In other words, the router is giving you double the bang for your bu…”
Anastasios Angelopoulos May 29, 2025 ▶ 1:30:12
Prediction Not checkable as stated
Chiang: Personalized leaderboards will improve LMArena data quality
“And in that case, we align the interests of individuals and the platform as a whole, because you don't want to mess up your personal leaderboard. Just like how people these days, when they use social media, They don't like a random post because if they do that…”
Wei-Lin Chiang May 29, 2025 ▶ 1:32:22
Disclosure
Chiang: LMArena open-sources all code, infrastructure, and models
“All the code infrastructures that we process the data is published as open source and also research blog, paper, and then including prompt leaderboard, we publish the paper. Open source, the models, the code, and everything.”
Wei-Lin Chiang May 29, 2025 ▶ 1:35:01
Prediction Held up
Angelopoulos: LMArena will remain open-source as a commercial company
“We're going to keep publishing papers. We're going to keep releasing open source. We're going to keep releasing open data.”
Anastasios Angelopoulos May 29, 2025 ▶ 1:36:21
Insight
Angelopoulos: Top AI researchers avoid companies building purely proprietary technology
“The best people don't want to hole up at a company and develop a bunch of proprietary technology that, you know, is never going to be released.”
Anastasios Angelopoulos May 29, 2025 ▶ 1:36:34
Assertion Supported
Chiang: Red Team Arena maintains a leaderboard for AI jailbreakers
“So in Red Team Arena, we have a leaderboard, not just for model, but for user, for job breakers, who is the best job breakers that can like identify issues. For all different models.”
Wei-Lin Chiang May 29, 2025 ▶ 1:40:59
Insight
Angelopoulos: High refusal rates do not make AI models inherently superior
“It's not necessarily the model that's like most, like refuses the most to answer these like queries that people ask necessarily better. Some people want a model that's more controllable. Some people want a model that's going to say whatever they want. Some peo…”
Anastasios Angelopoulos May 29, 2025 ▶ 1:42:49
Prediction Not checkable as stated
Angelopoulos: Real-world testing will remain fundamental for evaluating AI agents
“The fundamental is organic, real-world testing with feedback. That's not going to change. I can tell you that that is not going to change.”
Anastasios Angelopoulos May 29, 2025 ▶ 1:44:12
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.