The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 7 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Assertion Supported
Arena sampled open-source models at 60/40, debunking Leaderboard Illusion paper
“But, you know, there, for example said that we were, that we only sampled, like, nine percent open source models and, like, you know, 60%, like, closed source models, and this created a gap between open and closed source. But in reality, we're actually really …”
Anastasios Angelopoulos Dec 31, 2025 ▶ 11:51 [State of Evals] LMArena's $1.7B Vision — Anastasios Angelopoulos, LMArena
Assertion Supported
Angelopoulos: OpenAI o1 crushed Chatbot Arena, proving the benchmark isn't saturated
“So there's this model and it crushed the benchmark. You know, it's just like really like a big gap. And what that's telling us is that it's not saturated yet. And so it's still measuring some signal that was encouraging point.”
Anastasios Angelopoulos Nov 1, 2024 ▶ 27:20 In the Arena: How LMSys changed LLM Benchmarking Forever
Assertion Supported
Angelopoulos: The Chatbot Arena leaderboard is currently not an apples-to-apples comparison
“None of the leaderboard currently is apples to apples, because you have, like, Gemini Flash, you have, you know, all sorts of tiny models, like Llama Like, eight B and four or five B are not apples to apples.”
Anastasios Angelopoulos Nov 1, 2024 ▶ 28:03 In the Arena: How LMSys changed LLM Benchmarking Forever
Assertion Supported
Angelopoulos: Arena received grants from Sequoia and a16z before incorporating
“He was not, you know, A-sixteen was not the only one to do this. We also had a great grant from Sequoia, but Ansh was in particular quite, quite supportive of us and, you know, gave us some resources in order to continue building out Arena before we even We're…”
Anastasios Angelopoulos Dec 31, 2025 ▶ 1:52 [State of Evals] LMArena's $1.7B Vision — Anastasios Angelopoulos, LMArena
Assertion Supported
Gradio scaled Arena to 1 million monthly active users before migration
“Gradio scaled us to a million Mal.”
Anastasios Angelopoulos Dec 31, 2025 ▶ 8:48 [State of Evals] LMArena's $1.7B Vision — Anastasios Angelopoulos, LMArena
Assertion Supported
Angelopoulos: Chatbot Arena scores are calculated via logistic regression
“The arena score that we show on our leaderboard is a particular type of linear model, right? It's a linear model that takes, it's a logistic regression that takes model identities and fits them against human preference, right? So it regresses human preference …”
Anastasios Angelopoulos Nov 1, 2024 ▶ 14:57 In the Arena: How LMSys changed LLM Benchmarking Forever
Assertion Partly supported
Angelopoulos: LMSYS controls for markdown and lists in Arena rankings
“We have, you know, five, six different style components that have to do with markdown headers and bulleted lists and so on that we add here.”
Anastasios Angelopoulos Nov 1, 2024 ▶ 16:24 In the Arena: How LMSys changed LLM Benchmarking Forever
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.