The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 12 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 1 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Assertion Not checkable as stated
Angelopoulos: Chatbot Arena is immune to model overfitting by design
“Static benchmarks overfit. Why? It's because as Jan said earlier, you're giving the student the same test over and over. You have a model, you test it, you know, you look at whether or not it's improved on a static data set. Then you find another model, you te…”
Anastasios Angelopoulos May 29, 2025 ▶ 21:00 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Assertion Supported
Angelopoulos: LMArena makes style control the default AI evaluation method
“That's why we're making style control default.”
Anastasios Angelopoulos May 29, 2025 ▶ 12:55 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Prediction Not checkable as stated
Angelopoulos: Future AI evaluation will shift to personalized user leaderboards
“Absolutely. Absolutely. It should be personalized just for you. You should understand which models are best for you.”
Anastasios Angelopoulos May 29, 2025 ▶ 26:06 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Assertion Supported
Angelopoulos: LMArena router model outperforms all constituent models on Chatbot Arena
“When you train a prompt to leaderboard model, which is like, let's say a seven billion parameter model, and then you use it to route on just questions on the arena and everybody's questions, that model does better than any of the constituent models that were u…”
Anastasios Angelopoulos May 29, 2025 ▶ 1:29:10 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Assertion Supported
Angelopoulos: LMArena prompt router yields double the performance per dollar
“Now, if you trace the performance, the best performance that, you know, any individual model can give you as part of the router as a function of cost. That's like two X worse than the router. In other words, the router is giving you double the bang for your bu…”
Anastasios Angelopoulos May 29, 2025 ▶ 1:30:12 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Assertion Open · timeframe May 2026
Angelopoulos: Human evaluators prefer longer AI responses given equal content
“It's true that people vote for longer responses, you know, preferentially over shorter responses, even given the same contents or well-known human bias.”
Anastasios Angelopoulos May 29, 2025 ▶ 12:31 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Assertion Not checkable as stated
Angelopoulos: LMArena measures user preference, not AGI progress
“We don't claim to be an AGI benchmark. We are faithfully representing the preferences of our community.”
Anastasios Angelopoulos May 29, 2025 ▶ 18:08 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Assertion Supported
Angelopoulos: Bradley-Terry models converge for AI evaluation, unlike Elo scores
“Okay, let's move from Elo to Bradley Terry because we're actually performing an estimate here instead of just like You know, and the ELO score moves over time. It doesn't converge, but Rally Terry models converge and how do we then construct confidence interva…”
Anastasios Angelopoulos May 29, 2025 ▶ 38:49 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Assertion Not checkable as stated
Angelopoulos: Chatbot Arena has 1M+ monthly users and 150M+ conversations
“A lot of people don't know this, but ShopBot Arena is Used by like a million plus monthly users. We get like, you know, tens of thousands of votes on a daily basis. We have like over like, you know, a hundred fifty million conversations that have been had on t…”
Anastasios Angelopoulos May 29, 2025 ▶ 1:13:19 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Prediction Held up
Angelopoulos: LMArena will remain open-source as a commercial company
“We're going to keep publishing papers. We're going to keep releasing open source. We're going to keep releasing open data.”
Anastasios Angelopoulos May 29, 2025 ▶ 1:36:21 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Prediction Not checkable as stated
Angelopoulos: Real-world testing will remain fundamental for evaluating AI agents
“The fundamental is organic, real-world testing with feedback. That's not going to change. I can tell you that that is not going to change.”
Anastasios Angelopoulos May 29, 2025 ▶ 1:44:12 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Prediction Didn’t hold up
Angelopoulos: LMArena will launch Data-Driven Debugging within months
“So we're building a project now that we call data-driven debugging D three. It's, you know, it's a little farther out. It'll come in a couple months.”
Anastasios Angelopoulos May 29, 2025 ▶ 1:26:14 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.