Dec 31, 2025 · 24m · latent-space

[State of Evals] LMArena's $1.7B Vision — Anastasios Angelopoulos, LMArena

Anastasios Angelopoulos · 14m spoken Shawn Wang · 5m spoken Alessio Fanelli · 17s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this interview, Arena co-founder Anastasios Angelopoulos discusses the platform's evolution from a Berkeley research initiative into a venture-backed evaluation leader, detailing its organic benchmarking methodology, technical scaling, and multimodal expansion. He explains how Arena maintains platform integrity and subsidizes millions of blind model battles to establish an impartial North Star for frontier AI evaluation.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 29.2% of the talking time here. How this is scored →

The hosts as informed peer 4.8 Guest teaching 4.0 Guest disagreement 2.2 The hosts pushing back 2.8
05100:0010:0020:000:07–3:35 · The hosts as informed peer 3/10 Rebranding LMSYS to ARENA and Broadening Scope The conversation starts warmly with light banter about branding, dropping 'LM' from LMSYS, and Ansh Sharma's early incubation of the company. Swyx and Alessio ask open-ended questions about the transition from Berkeley research project to a venture-backed startup.3:36–6:01 · The hosts as informed peer 5/10 Deploying $100M and Measuring Organic User Scale Swyx presses Anastasios on how ARENA will deploy $100M, questioning whether they pay enterprise inference rates and how they verify organic user distributions without mandatory logins. Anastasios explains their cost structure, user base scale, and survey methodology.6:01–8:36 · The hosts as informed peer 5/10 Arena Methodology vs. Competitors and Prompt Engineering The hosts bring up competing benchmark platforms like Artificial Analysis. Swyx defends the utility of pre-rendered prompts for users learning how to prompt, while Anastasios emphasizes ARENA's organic, user-driven query methodology.8:36–13:38 · The hosts as informed peer 6/10 Technical Infrastructure: Migrating from Gradio to React After briefly touching on migrating from Gradio to React, Swyx raises Cohere's 'Leaderboard Illusion' paper. Anastasios strongly critiques the paper as unscientific and refutes its claims regarding open-source sampling ratios and pre-release testing.13:38–19:07 · The hosts as informed peer 5/10 The Economic and Practical Value of Multimodal Generation Swyx discusses his initial skepticism toward image generation before highlighting real-world workflows with Nano Banana Pro. Anastasios outlines ARENA's non-negotiable integrity standards, insisting the public leaderboard operates strictly as an impartial loss leader.19:08–21:44 · The hosts as informed peer 5/10 Strategic Product Focus, Community Growth, and User Retention The discussion turns to product roadmap, resisting API feature creep, and consumer retention dynamics. Swyx pushes on the need to evaluate agentic software harnesses like Devin rather than raw foundational models alone.0:07–3:35 · Guest teaching 2/10 Rebranding LMSYS to ARENA and Broadening Scope The conversation starts warmly with light banter about branding, dropping 'LM' from LMSYS, and Ansh Sharma's early incubation of the company. Swyx and Alessio ask open-ended questions about the transition from Berkeley research project to a venture-backed startup.3:36–6:01 · Guest teaching 4/10 Deploying $100M and Measuring Organic User Scale Swyx presses Anastasios on how ARENA will deploy $100M, questioning whether they pay enterprise inference rates and how they verify organic user distributions without mandatory logins. Anastasios explains their cost structure, user base scale, and survey methodology.6:01–8:36 · Guest teaching 4/10 Arena Methodology vs. Competitors and Prompt Engineering The hosts bring up competing benchmark platforms like Artificial Analysis. Swyx defends the utility of pre-rendered prompts for users learning how to prompt, while Anastasios emphasizes ARENA's organic, user-driven query methodology.8:36–13:38 · Guest teaching 6/10 Technical Infrastructure: Migrating from Gradio to React After briefly touching on migrating from Gradio to React, Swyx raises Cohere's 'Leaderboard Illusion' paper. Anastasios strongly critiques the paper as unscientific and refutes its claims regarding open-source sampling ratios and pre-release testing.13:38–19:07 · Guest teaching 4/10 The Economic and Practical Value of Multimodal Generation Swyx discusses his initial skepticism toward image generation before highlighting real-world workflows with Nano Banana Pro. Anastasios outlines ARENA's non-negotiable integrity standards, insisting the public leaderboard operates strictly as an impartial loss leader.19:08–21:44 · Guest teaching 4/10 Strategic Product Focus, Community Growth, and User Retention The discussion turns to product roadmap, resisting API feature creep, and consumer retention dynamics. Swyx pushes on the need to evaluate agentic software harnesses like Devin rather than raw foundational models alone.0:07–3:35 · Guest disagreement 1/10 Rebranding LMSYS to ARENA and Broadening Scope The conversation starts warmly with light banter about branding, dropping 'LM' from LMSYS, and Ansh Sharma's early incubation of the company. Swyx and Alessio ask open-ended questions about the transition from Berkeley research project to a venture-backed startup.3:36–6:01 · Guest disagreement 2/10 Deploying $100M and Measuring Organic User Scale Swyx presses Anastasios on how ARENA will deploy $100M, questioning whether they pay enterprise inference rates and how they verify organic user distributions without mandatory logins. Anastasios explains their cost structure, user base scale, and survey methodology.6:01–8:36 · Guest disagreement 2/10 Arena Methodology vs. Competitors and Prompt Engineering The hosts bring up competing benchmark platforms like Artificial Analysis. Swyx defends the utility of pre-rendered prompts for users learning how to prompt, while Anastasios emphasizes ARENA's organic, user-driven query methodology.8:36–13:38 · Guest disagreement 4/10 Technical Infrastructure: Migrating from Gradio to React After briefly touching on migrating from Gradio to React, Swyx raises Cohere's 'Leaderboard Illusion' paper. Anastasios strongly critiques the paper as unscientific and refutes its claims regarding open-source sampling ratios and pre-release testing.13:38–19:07 · Guest disagreement 2/10 The Economic and Practical Value of Multimodal Generation Swyx discusses his initial skepticism toward image generation before highlighting real-world workflows with Nano Banana Pro. Anastasios outlines ARENA's non-negotiable integrity standards, insisting the public leaderboard operates strictly as an impartial loss leader.19:08–21:44 · Guest disagreement 2/10 Strategic Product Focus, Community Growth, and User Retention The discussion turns to product roadmap, resisting API feature creep, and consumer retention dynamics. Swyx pushes on the need to evaluate agentic software harnesses like Devin rather than raw foundational models alone.0:07–3:35 · The hosts pushing back 1/10 Rebranding LMSYS to ARENA and Broadening Scope The conversation starts warmly with light banter about branding, dropping 'LM' from LMSYS, and Ansh Sharma's early incubation of the company. Swyx and Alessio ask open-ended questions about the transition from Berkeley research project to a venture-backed startup.3:36–6:01 · The hosts pushing back 4/10 Deploying $100M and Measuring Organic User Scale Swyx presses Anastasios on how ARENA will deploy $100M, questioning whether they pay enterprise inference rates and how they verify organic user distributions without mandatory logins. Anastasios explains their cost structure, user base scale, and survey methodology.6:01–8:36 · The hosts pushing back 4/10 Arena Methodology vs. Competitors and Prompt Engineering The hosts bring up competing benchmark platforms like Artificial Analysis. Swyx defends the utility of pre-rendered prompts for users learning how to prompt, while Anastasios emphasizes ARENA's organic, user-driven query methodology.8:36–13:38 · The hosts pushing back 2/10 Technical Infrastructure: Migrating from Gradio to React After briefly touching on migrating from Gradio to React, Swyx raises Cohere's 'Leaderboard Illusion' paper. Anastasios strongly critiques the paper as unscientific and refutes its claims regarding open-source sampling ratios and pre-release testing.13:38–19:07 · The hosts pushing back 3/10 The Economic and Practical Value of Multimodal Generation Swyx discusses his initial skepticism toward image generation before highlighting real-world workflows with Nano Banana Pro. Anastasios outlines ARENA's non-negotiable integrity standards, insisting the public leaderboard operates strictly as an impartial loss leader.19:08–21:44 · The hosts pushing back 3/10 Strategic Product Focus, Community Growth, and User Retention The discussion turns to product roadmap, resisting API feature creep, and consumer retention dynamics. Swyx pushes on the need to evaluate agentic software harnesses like Devin rather than raw foundational models alone.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 38.5% · guest 61.5%0:00 · the hosts 38.5% · guest 61.5%3:00 · the hosts 16.6% · guest 83.4%3:00 · the hosts 16.6% · guest 83.4%6:00 · the hosts 34.9% · guest 65.1%6:00 · the hosts 34.9% · guest 65.1%9:00 · the hosts 22.1% · guest 77.9%9:00 · the hosts 22.1% · guest 77.9%12:00 · the hosts 45.6% · guest 54.4%12:00 · the hosts 45.6% · guest 54.4%15:00 · the hosts 39.2% · guest 60.8%15:00 · the hosts 39.2% · guest 60.8%18:00 · the hosts 12.5% · guest 87.5%18:00 · the hosts 12.5% · guest 87.5%21:00 · the hosts 24.3% · guest 75.7%21:00 · the hosts 24.3% · guest 75.7%24:00 · the hosts 0% · guest 0%24:00 · the hosts 0% · guest 0%
Sharpest disagreement ▶ 10:30 Dismissing the Leaderboard Illusion paper as unscientific

Anastasios directly dismisses the Cohere research paper critiquing LMSYS as fundamentally flawed and unscientific, challenging the motivations and competence behind its claims.

Hardest push from the hosts ▶ 8:03 Swyx defending pre-generated prompt evaluations

Swyx challenges Anastasios's dismissive stance on pre-rendered video arenas by arguing that users genuinely need to see outside examples to learn effective prompting.

Biggest teaching moment ▶ 11:40 Dissecting the data errors in Cohere's critique

Anastasios educates the hosts on the exact factual errors published in the critique, explaining the true 60/40 open-to-closed sampling distribution and the community purpose of blind preview testing.

The host holds their own ▶ 15:04 Swyx showcasing multimodal paper diagram generation

Swyx demonstrates his deep technical and content workflow expertise by detailing how he feeds dense RL research papers into Nano Banana Pro to produce PhD-level diagrams instantly.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Rebranding LMSYS to ARENA and Broadening Scope 3211 The conversation starts warmly with light banter about branding, dropping 'LM' from LMSYS, and Ansh Sharma's early incubation of the company. Swyx and Alessio ask open-ended questions about the transition from Berkeley research project to a venture-backed startup.
Deploying $100M and Measuring Organic User Scale 5424 Swyx presses Anastasios on how ARENA will deploy $100M, questioning whether they pay enterprise inference rates and how they verify organic user distributions without mandatory logins. Anastasios explains their cost structure, user base scale, and survey methodology.
Arena Methodology vs. Competitors and Prompt Engineering 5424 The hosts bring up competing benchmark platforms like Artificial Analysis. Swyx defends the utility of pre-rendered prompts for users learning how to prompt, while Anastasios emphasizes ARENA's organic, user-driven query methodology.
Technical Infrastructure: Migrating from Gradio to React 6642 After briefly touching on migrating from Gradio to React, Swyx raises Cohere's 'Leaderboard Illusion' paper. Anastasios strongly critiques the paper as unscientific and refutes its claims regarding open-source sampling ratios and pre-release testing.
The Economic and Practical Value of Multimodal Generation 5423 Swyx discusses his initial skepticism toward image generation before highlighting real-world workflows with Nano Banana Pro. Anastasios outlines ARENA's non-negotiable integrity standards, insisting the public leaderboard operates strictly as an impartial loss leader.
Strategic Product Focus, Community Growth, and User Retention 5423 The discussion turns to product roadmap, resisting API feature creep, and consumer retention dynamics. Swyx pushes on the need to evaluate agentic software harnesses like Devin rather than raw foundational models alone.

Statements from this episode (18)

Disclosure
Angelopoulos: Arena Dropped 'LM' to Broaden Beyond Language Models
“So, so we wanted to maybe broaden a little bit. And we were the first Serena, so we feel like let's kind of try to own that.”
Anastasios Angelopoulos Dec 31, 2025 ▶ 0:57
Assertion Supported
Angelopoulos: Arena received grants from Sequoia and a16z before incorporating
“He was not, you know, A-sixteen was not the only one to do this. We also had a great grant from Sequoia, but Ansh was in particular quite, quite supportive of us and, you know, gave us some resources in order to continue building out Arena before we even We're…”
Anastasios Angelopoulos Dec 31, 2025 ▶ 1:52
Assertion Not checkable as stated
LMArena funds all model inference running on its platform
“We fund all of the inference on the platform.”
Anastasios Angelopoulos Dec 31, 2025 ▶ 4:13
Disclosure
Angelopoulos: LMArena receives only standard enterprise inference discounts
“No, no, we get discounts, but they're, but they are standard enterprise discounts. The same that would be given to any other customer.”
Anastasios Angelopoulos Dec 31, 2025 ▶ 4:19
Assertion Not checkable as stated
Arena processes tens of millions of conversations monthly, totaling 250 million
“We have probably two hundred and fifty million conversations that happen over the course of the platform. We're on the order of, you know, mid tens of millions of conversations every month that are happening on the platform.”
Anastasios Angelopoulos Dec 31, 2025 ▶ 4:46
Assertion Not checkable as stated
Angelopoulos: 25% of LMArena platform users write software for a living
“25% of the people on our platform, for example, do software for a living.”
Anastasios Angelopoulos Dec 31, 2025 ▶ 5:12
Assertion Not checkable as stated
Angelopoulos: About half of LMArena users are authenticated
“About half of our users now are login.”
Anastasios Angelopoulos Dec 31, 2025 ▶ 5:34
Opinion
Arena's organic user prompts provide realism that Artificial Analysis lacks
“They have arenas, but the arenas are not based on organic usage. Like the thing that distinguishes our platform versus theirs is that the users are actually inputting their own use case. They're actually asking their own question. And that gives a level of rea…”
Anastasios Angelopoulos Dec 31, 2025 ▶ 7:24
Assertion Supported
swyx: Artificial Analysis video arena uses pre-generated videos instead of user inputs
“So like, for example, for AA, their video arena is pre-generated videos. You can't enter in your own video.”
Shawn Wang Dec 31, 2025 ▶ 7:50
Assertion Supported
Gradio scaled Arena to 1 million monthly active users before migration
“Gradio scaled us to a million Mal.”
Anastasios Angelopoulos Dec 31, 2025 ▶ 8:48
Disclosure
Angelopoulos: LMArena's top expenses are free-tier inference, hiring, and SF office
“Primarily inference that funds the free usage of the platform and then also hiring, of course, headcount. We have an office, you know. That's an SF.”
Anastasios Angelopoulos Dec 31, 2025 ▶ 9:57
Assertion Supported
Arena sampled open-source models at 60/40, debunking Leaderboard Illusion paper
“But, you know, there, for example said that we were, that we only sampled, like, nine percent open source models and, like, you know, 60%, like, closed source models, and this created a gap between open and closed source. But in reality, we're actually really …”
Anastasios Angelopoulos Dec 31, 2025 ▶ 11:51
Assertion Not checkable as stated
Arena's anonymous Nano Banana test moved Google's stock and product roadmap
“I mean, that moment alone changed Google's like roadmap. Market share. Seriously. I mean, Google stock, billions of dollars are moving because of Nano.”
Anastasios Angelopoulos Dec 31, 2025 ▶ 13:20
Prediction Not checkable as stated
Angelopoulos: Academic paper figures will soon be generated by AI models
“Soon we're not going to be even making them for our papers. We're, they're just going to be, our paper figures are going to be made by Emily.”
Anastasios Angelopoulos Dec 31, 2025 ▶ 14:59
Assertion Not checkable as stated
Angelopoulos: LMArena has released more real-world AI data than almost anyone
“We've probably released more data than basically anybody on the real world use cases of AI.”
Anastasios Angelopoulos Dec 31, 2025 ▶ 16:32
Disclosure
Arena's public leaderboard will never adopt Gartner-style pay-to-play models
“You can't pay to get on the public leaderboard. It's not like a Gartner in that sense. It's not like any of these, like you know, pay to play systems, never going to be like that. Models are going to be listed on the leaderboard, whether or not the providers p…”
Anastasios Angelopoulos Dec 31, 2025 ▶ 17:26
Disclosure
Angelopoulos: LMArena to launch video evaluations by early next year
“Video we're soon to launch on the site at some point, you know, later this year or early next.”
Anastasios Angelopoulos Dec 31, 2025 ▶ 18:58
Opinion
Angelopoulos rejects claims that Cognition's Devin is dead
“Devin's not gone. Devin's everywhere.”
Anastasios Angelopoulos Dec 31, 2025 ▶ 23:20
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.