Dec 31, 2025 · 24m · latent-space
[State of Evals] LMArena's $1.7B Vision — Anastasios Angelopoulos, LMArena
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this interview, Arena co-founder Anastasios Angelopoulos discusses the platform's evolution from a Berkeley research initiative into a venture-backed evaluation leader, detailing its organic benchmarking methodology, technical scaling, and multimodal expansion. He explains how Arena maintains platform integrity and subsidizes millions of blind model battles to establish an impartial North Star for frontier AI evaluation.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 29.2% of the talking time here. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Anastasios directly dismisses the Cohere research paper critiquing LMSYS as fundamentally flawed and unscientific, challenging the motivations and competence behind its claims.
Hardest push from the hosts ▶ 8:03 Swyx defending pre-generated prompt evaluationsSwyx challenges Anastasios's dismissive stance on pre-rendered video arenas by arguing that users genuinely need to see outside examples to learn effective prompting.
Biggest teaching moment ▶ 11:40 Dissecting the data errors in Cohere's critiqueAnastasios educates the hosts on the exact factual errors published in the critique, explaining the true 60/40 open-to-closed sampling distribution and the community purpose of blind preview testing.
The host holds their own ▶ 15:04 Swyx showcasing multimodal paper diagram generationSwyx demonstrates his deep technical and content workflow expertise by detailing how he feeds dense RL research papers into Nano Banana Pro to produce PhD-level diagrams instantly.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Rebranding LMSYS to ARENA and Broadening Scope | 3 | 2 | 1 | 1 | The conversation starts warmly with light banter about branding, dropping 'LM' from LMSYS, and Ansh Sharma's early incubation of the company. Swyx and Alessio ask open-ended questions about the transition from Berkeley research project to a venture-backed startup. | |
| Deploying $100M and Measuring Organic User Scale | 5 | 4 | 2 | 4 | Swyx presses Anastasios on how ARENA will deploy $100M, questioning whether they pay enterprise inference rates and how they verify organic user distributions without mandatory logins. Anastasios explains their cost structure, user base scale, and survey methodology. | |
| Arena Methodology vs. Competitors and Prompt Engineering | 5 | 4 | 2 | 4 | The hosts bring up competing benchmark platforms like Artificial Analysis. Swyx defends the utility of pre-rendered prompts for users learning how to prompt, while Anastasios emphasizes ARENA's organic, user-driven query methodology. | |
| Technical Infrastructure: Migrating from Gradio to React | 6 | 6 | 4 | 2 | After briefly touching on migrating from Gradio to React, Swyx raises Cohere's 'Leaderboard Illusion' paper. Anastasios strongly critiques the paper as unscientific and refutes its claims regarding open-source sampling ratios and pre-release testing. | |
| The Economic and Practical Value of Multimodal Generation | 5 | 4 | 2 | 3 | Swyx discusses his initial skepticism toward image generation before highlighting real-world workflows with Nano Banana Pro. Anastasios outlines ARENA's non-negotiable integrity standards, insisting the public leaderboard operates strictly as an impartial loss leader. | |
| Strategic Product Focus, Community Growth, and User Retention | 5 | 4 | 2 | 3 | The discussion turns to product roadmap, resisting API feature creep, and consumer retention dynamics. Swyx pushes on the need to evaluate agentic software harnesses like Devin rather than raw foundational models alone. |