The Ledger, every show

Every statement that passed quotation and attribution checks, across all 44 shows. Pick shows below, then mix any filter with any other.

shows every show 44 of 44
every show
clear all ✕
a16z Assertion Supported
Angelopoulos: LMArena prompt router yields double the performance per dollar
“Now, if you trace the performance, the best performance that, you know, any individual model can give you as part of the router as a function of cost. That's like two X worse than the router. In other words, the router is giving you double the bang for your bu…”
Anastasios Angelopoulos May 29, 2025 ▶ 1:30:12 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
a16z Insight
Angelopoulos: Top AI researchers avoid companies building purely proprietary technology
“The best people don't want to hole up at a company and develop a bunch of proprietary technology that, you know, is never going to be released.”
Anastasios Angelopoulos May 29, 2025 ▶ 1:36:34 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
a16z Insight
Angelopoulos: High refusal rates do not make AI models inherently superior
“It's not necessarily the model that's like most, like refuses the most to answer these like queries that people ask necessarily better. Some people want a model that's more controllable. Some people want a model that's going to say whatever they want. Some peo…”
Anastasios Angelopoulos May 29, 2025 ▶ 1:42:49 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
LATENT SPACE Assertion Supported
Angelopoulos: OpenAI o1 crushed Chatbot Arena, proving the benchmark isn't saturated
“So there's this model and it crushed the benchmark. You know, it's just like really like a big gap. And what that's telling us is that it's not saturated yet. And so it's still measuring some signal that was encouraging point.”
Anastasios Angelopoulos Nov 1, 2024 ▶ 27:20 In the Arena: How LMSys changed LLM Benchmarking Forever
LATENT SPACE Assertion Supported
Angelopoulos: The Chatbot Arena leaderboard is currently not an apples-to-apples comparison
“None of the leaderboard currently is apples to apples, because you have, like, Gemini Flash, you have, you know, all sorts of tiny models, like Llama Like, eight B and four or five B are not apples to apples.”
Anastasios Angelopoulos Nov 1, 2024 ▶ 28:03 In the Arena: How LMSys changed LLM Benchmarking Forever
Angelopoulos: Highly effective LLM routers only need simple heuristics like length
“Well, I think that you can build a very, very simple router that is very effective. So let me give you an example. You can build a great router with one parameter, and the parameter is just like, I'm gonna check if my question is hard, and if it's hard, then I…”
Anastasios Angelopoulos Nov 1, 2024 ▶ 36:33 In the Arena: How LMSys changed LLM Benchmarking Forever
20VC Assertion Supported
Angelopoulos: Chinese AI labs face severe hardware constraints and rely on black-market chips
“So they're way hardware constrained over there. And they've been trying to like black market import chips because of this. And you see this in the news, right? The information just reported on this.”
Anastasios Angelopoulos Aug 2, 2026 ▶ 14:01 Arena CEO: There Will be a $100BN US Open-Source Model & Data is a Trillion Dollar Market
20VC Assertion Partly supported
Angelopoulos: China Has Already Restricted American AI Models Domestically
“And by the way, it's worth noting that China has already restricted the use of American models within China, right? So if you look at the two by two matrix of US China restrict, not restrict, you know, like export import stuff They have already restricted the …”
Anastasios Angelopoulos Aug 2, 2026 ▶ 16:10 Arena CEO: There Will be a $100BN US Open-Source Model & Data is a Trillion Dollar Market
20VC Insight
Angelopoulos: Open-source AI growth reduces Nvidia's revenue concentration
“Of course, Jensen is in some sense self-serving with this letter, because the more open source models are developed, the more companies are going to be training on GPUs. They're going to be fine tuning on their own data. And it's just more and more spend. It d…”
Anastasios Angelopoulos Aug 2, 2026 ▶ 21:17 Arena CEO: There Will be a $100BN US Open-Source Model & Data is a Trillion Dollar Market
20VC Disclosure
Angelopoulos reveals Arena uses Alibaba's Qwen in its tech stack
“And I said, you know, yeah, we use Quinn for X, Y, Z.”
Anastasios Angelopoulos Aug 2, 2026 ▶ 23:38 Arena CEO: There Will be a $100BN US Open-Source Model & Data is a Trillion Dollar Market
20VC Assertion Supported
Angelopoulos: Google's Gemma sits on the performance-versus-cost Pareto curve
“Gemma, by the way, is pretty good in terms of efficiency. If you look at arena, you'll see the, on the Pareto curves of like performance versus cost. Gemma's on there.”
Anastasios Angelopoulos Aug 2, 2026 ▶ 24:37 Arena CEO: There Will be a $100BN US Open-Source Model & Data is a Trillion Dollar Market
20VC Assertion Supported
Angelopoulos: Harvey's CEO views model labs as his biggest competitive worry
“Harvey, the CEO of Harvey himself is saying that, you know, his biggest competitive worry is the model labs.”
Anastasios Angelopoulos Aug 2, 2026 ▶ 56:44 Arena CEO: There Will be a $100BN US Open-Source Model & Data is a Trillion Dollar Market
LATENT SPACE Disclosure
Angelopoulos: LMArena receives only standard enterprise inference discounts
“No, no, we get discounts, but they're, but they are standard enterprise discounts. The same that would be given to any other customer.”
Anastasios Angelopoulos Dec 31, 2025 ▶ 4:19 [State of Evals] LMArena's $1.7B Vision — Anastasios Angelopoulos, LMArena
LATENT SPACE Assertion Not checkable as stated
Arena processes tens of millions of conversations monthly, totaling 250 million
“We have probably two hundred and fifty million conversations that happen over the course of the platform. We're on the order of, you know, mid tens of millions of conversations every month that are happening on the platform.”
Anastasios Angelopoulos Dec 31, 2025 ▶ 4:46 [State of Evals] LMArena's $1.7B Vision — Anastasios Angelopoulos, LMArena
LATENT SPACE Prediction Not checkable as stated
Angelopoulos: Academic paper figures will soon be generated by AI models
“Soon we're not going to be even making them for our papers. We're, they're just going to be, our paper figures are going to be made by Emily.”
Anastasios Angelopoulos Dec 31, 2025 ▶ 14:59 [State of Evals] LMArena's $1.7B Vision — Anastasios Angelopoulos, LMArena
LATENT SPACE Assertion Not checkable as stated
Angelopoulos: LMArena has released more real-world AI data than almost anyone
“We've probably released more data than basically anybody on the real world use cases of AI.”
Anastasios Angelopoulos Dec 31, 2025 ▶ 16:32 [State of Evals] LMArena's $1.7B Vision — Anastasios Angelopoulos, LMArena
a16z Assertion Open · timeframe May 2026
Angelopoulos: Human evaluators prefer longer AI responses given equal content
“It's true that people vote for longer responses, you know, preferentially over shorter responses, even given the same contents or well-known human bias.”
Anastasios Angelopoulos May 29, 2025 ▶ 12:31 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
a16z Assertion Not checkable as stated
Angelopoulos: LMArena measures user preference, not AGI progress
“We don't claim to be an AGI benchmark. We are faithfully representing the preferences of our community.”
Anastasios Angelopoulos May 29, 2025 ▶ 18:08 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
a16z Disclosure
Angelopoulos: LMArena is building personalized AI leaderboard tools
“Yeah, and we should be giving you the tools to do that, and we're currently building them.”
Anastasios Angelopoulos May 29, 2025 ▶ 25:56 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
a16z Assertion Supported
Angelopoulos: Bradley-Terry models converge for AI evaluation, unlike Elo scores
“Okay, let's move from Elo to Bradley Terry because we're actually performing an estimate here instead of just like You know, and the ELO score moves over time. It doesn't converge, but Rally Terry models converge and how do we then construct confidence interva…”
Anastasios Angelopoulos May 29, 2025 ▶ 38:49 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
a16z Insight
Angelopoulos: Reinforcement learning allows AI models to surpass human teachers
“And supervised learning, you can only do as well as the best human that you have. Because what's happening is that you're learning from the teacher. In reinforcement learning, you're learning from the world. You're able to learn things better than the best hum…”
Anastasios Angelopoulos May 29, 2025 ▶ 1:00:48 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
a16z Assertion Not checkable as stated
Angelopoulos: Chatbot Arena has 1M+ monthly users and 150M+ conversations
“A lot of people don't know this, but ShopBot Arena is Used by like a million plus monthly users. We get like, you know, tens of thousands of votes on a daily basis. We have like over like, you know, a hundred fifty million conversations that have been had on t…”
Anastasios Angelopoulos May 29, 2025 ▶ 1:13:19 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
a16z Prediction Held up
Angelopoulos: LMArena will remain open-source as a commercial company
“We're going to keep publishing papers. We're going to keep releasing open source. We're going to keep releasing open data.”
Anastasios Angelopoulos May 29, 2025 ▶ 1:36:21 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
a16z Prediction Not checkable as stated
Angelopoulos: Real-world testing will remain fundamental for evaluating AI agents
“The fundamental is organic, real-world testing with feedback. That's not going to change. I can tell you that that is not going to change.”
Anastasios Angelopoulos May 29, 2025 ▶ 1:44:12 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
LATENT SPACE Disclosure
Angelopoulos: LMSYS considers default style control but avoids imposing opinions
“We consider that we're still actively considering it. It's just, you know, once you make that step, once you take that step, you're introducing your opinion. And I'm not, you know, why should our opinion be the one? That's kind of a community choice. We could …”
Anastasios Angelopoulos Nov 1, 2024 ▶ 18:27 In the Arena: How LMSys changed LLM Benchmarking Forever
LATENT SPACE Assertion Not checkable as stated
Angelopoulos: Five-model selection bias is tiny compared to voter variability
“We don't do that right now, partially because we kind of have know from simulations that the amount of selection bias you incur with these five things is just not huge. It's not huge in comparison to the variability that you get from the, from just regular hum…”
Anastasios Angelopoulos Nov 1, 2024 ▶ 31:47 In the Arena: How LMSys changed LLM Benchmarking Forever
Angelopoulos: Live voter data asymptotically eliminates pre-release ELO bias
“What happened is that over time, because we're getting new data, it'll get adjusted down. So if there's any bias that gets introduced at that stage in the long run, it actually doesn't matter because asymptotically, basically like in the long run, there's way …”
Anastasios Angelopoulos Nov 1, 2024 ▶ 32:16 In the Arena: How LMSys changed LLM Benchmarking Forever
LATENT SPACE Disclosure
Angelopoulos: Arena Dropped 'LM' to Broaden Beyond Language Models
“So, so we wanted to maybe broaden a little bit. And we were the first Serena, so we feel like let's kind of try to own that.”
Anastasios Angelopoulos Dec 31, 2025 ▶ 0:57 [State of Evals] LMArena's $1.7B Vision — Anastasios Angelopoulos, LMArena
LATENT SPACE Assertion Supported
Angelopoulos: Arena received grants from Sequoia and a16z before incorporating
“He was not, you know, A-sixteen was not the only one to do this. We also had a great grant from Sequoia, but Ansh was in particular quite, quite supportive of us and, you know, gave us some resources in order to continue building out Arena before we even We're…”
Anastasios Angelopoulos Dec 31, 2025 ▶ 1:52 [State of Evals] LMArena's $1.7B Vision — Anastasios Angelopoulos, LMArena
LATENT SPACE Assertion Not checkable as stated
LMArena funds all model inference running on its platform
“We fund all of the inference on the platform.”
Anastasios Angelopoulos Dec 31, 2025 ▶ 4:13 [State of Evals] LMArena's $1.7B Vision — Anastasios Angelopoulos, LMArena
LATENT SPACE Assertion Not checkable as stated
Angelopoulos: 25% of LMArena platform users write software for a living
“25% of the people on our platform, for example, do software for a living.”
Anastasios Angelopoulos Dec 31, 2025 ▶ 5:12 [State of Evals] LMArena's $1.7B Vision — Anastasios Angelopoulos, LMArena
LATENT SPACE Assertion Not checkable as stated
Angelopoulos: About half of LMArena users are authenticated
“About half of our users now are login.”
Anastasios Angelopoulos Dec 31, 2025 ▶ 5:34 [State of Evals] LMArena's $1.7B Vision — Anastasios Angelopoulos, LMArena
LATENT SPACE Assertion Supported
Gradio scaled Arena to 1 million monthly active users before migration
“Gradio scaled us to a million Mal.”
Anastasios Angelopoulos Dec 31, 2025 ▶ 8:48 [State of Evals] LMArena's $1.7B Vision — Anastasios Angelopoulos, LMArena
LATENT SPACE Disclosure
Angelopoulos: LMArena's top expenses are free-tier inference, hiring, and SF office
“Primarily inference that funds the free usage of the platform and then also hiring, of course, headcount. We have an office, you know. That's an SF.”
Anastasios Angelopoulos Dec 31, 2025 ▶ 9:57 [State of Evals] LMArena's $1.7B Vision — Anastasios Angelopoulos, LMArena
LATENT SPACE Disclosure
Angelopoulos: LMArena to launch video evaluations by early next year
“Video we're soon to launch on the site at some point, you know, later this year or early next.”
Anastasios Angelopoulos Dec 31, 2025 ▶ 18:58 [State of Evals] LMArena's $1.7B Vision — Anastasios Angelopoulos, LMArena
a16z Disclosure
Angelopoulos: LMArena conducts pre-release model testing for AI developers
“One of the things that we help everybody to do is pre-release testing of their models. Okay. So it's not just that, you know, we work together to evaluate the models are released, but we also try to be their release partners and say, Hey, can we help you guys …”
Anastasios Angelopoulos May 29, 2025 ▶ 4:38 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
a16z Prediction Didn’t hold up
Angelopoulos: LMArena will launch Data-Driven Debugging within months
“So we're building a project now that we call data-driven debugging D three. It's, you know, it's a little farther out. It'll come in a couple months.”
Anastasios Angelopoulos May 29, 2025 ▶ 1:26:14 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
LATENT SPACE Assertion Supported
Angelopoulos: Chatbot Arena scores are calculated via logistic regression
“The arena score that we show on our leaderboard is a particular type of linear model, right? It's a linear model that takes, it's a logistic regression that takes model identities and fits them against human preference, right? So it regresses human preference …”
Anastasios Angelopoulos Nov 1, 2024 ▶ 14:57 In the Arena: How LMSys changed LLM Benchmarking Forever
LATENT SPACE Assertion Partly supported
Angelopoulos: LMSYS controls for markdown and lists in Arena rankings
“We have, you know, five, six different style components that have to do with markdown headers and bulleted lists and so on that we add here.”
Anastasios Angelopoulos Nov 1, 2024 ▶ 16:24 In the Arena: How LMSys changed LLM Benchmarking Forever
LATENT SPACE Disclosure
Angelopoulos: LMSYS wants to integrate live code execution in Chatbot Arena
“For example, it'd be great if we could execute code within Arena. It'd be fantastic. We want to do it.”
Anastasios Angelopoulos Nov 1, 2024 ▶ 21:49 In the Arena: How LMSys changed LLM Benchmarking Forever
LATENT SPACE Disclosure
Angelopoulos: Chatbot Arena is decoupling from LMSYS as co-creators shift focus
“Sort of Chatbot Arena has, of course, like, kind of become its own thing, and Lianmin and Ying, who are, you know, created LMSYS, have kind of, like, moved on to working on SGLang, and now They're doing other projects that are sort of originating from LMSS. An…”
Anastasios Angelopoulos Nov 1, 2024 ▶ 38:23 In the Arena: How LMSys changed LLM Benchmarking Forever
Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.