The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 6 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Insight
Angelopoulos: Most mission-critical AI queries are subjective, not factual lookups
“In reality, even in such industries, the majority of questions that people ask are subjective. Okay. So the mythology that in hard sciences or in mission critical industries, people just have like cut and dried questions and they just need like a retrieval and…”
Anastasios Angelopoulos May 29, 2025 ▶ 2:50 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Insight
Angelopoulos: AI evaluation performance follows a data scaling law
“Because language models are sort of the intermediary that gets you to this evaluation, there's also a scaling law that comes along with it. Which is to say that the more data you get, the bigger you build the platform, the better you can make your evaluations,…”
Anastasios Angelopoulos May 29, 2025 ▶ 56:18 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Insight
Angelopoulos: AI leaderboards can utilize any form of interaction feedback
“Pairwise comparison feedback is not the only kind of feedback that we can use to construct leaderboards. We can construct leaderboards with any form of feedback.”
Anastasios Angelopoulos May 29, 2025 ▶ 1:26:25 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Insight
Angelopoulos: Top AI researchers avoid companies building purely proprietary technology
“The best people don't want to hole up at a company and develop a bunch of proprietary technology that, you know, is never going to be released.”
Anastasios Angelopoulos May 29, 2025 ▶ 1:36:34 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Insight
Angelopoulos: High refusal rates do not make AI models inherently superior
“It's not necessarily the model that's like most, like refuses the most to answer these like queries that people ask necessarily better. Some people want a model that's more controllable. Some people want a model that's going to say whatever they want. Some peo…”
Anastasios Angelopoulos May 29, 2025 ▶ 1:42:49 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Insight
Angelopoulos: Reinforcement learning allows AI models to surpass human teachers
“And supervised learning, you can only do as well as the best human that you have. Because what's happening is that you're learning from the teacher. In reinforcement learning, you're learning from the world. You're able to learn things better than the best hum…”
Anastasios Angelopoulos May 29, 2025 ▶ 1:00:48 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.