Anastasios Angelopoulos

Co-Founder and CEO, Arena · 1 appearance on the record.

computed by AI from the episodes · how this works → · full disclaimer →

founderexecutivescientistauthor@ml_angelopoulos ↗LinkedIn ↗angelopoulos.ai ↗

He co-founded Arena, the crowdsourced human-preference benchmarking platform for AI foundation models originating out of UC Berkeley. A former UC Berkeley postdoctoral researcher, he specializes in reliable AI and conformal prediction.

22statements → 12claims → 6claims resolved → 83%fully supported → 4.5/5average certainty → 2.14/5average debate potential →

5 supported 0 partly supported 1 contradicted 1 not yet assessed 5 not checkable as stated how the 12 claims stand · each chip opens the sources

4 predictions · 8 assertions · 2 opinions · 6 insights · 2 disclosures · every statement was checked. The predictions and assertions are the 12 claims: statements the public record can support or contradict. 6 are resolved, 1 is not yet assessed, and 5 name no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Anastasios argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
Angelopoulos: LMArena makes style control the default AI evaluation method
“That's why we're making style control default.”
Anastasios Angelopoulos May 29, 2025 ▶ 12:55 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable

Their most notable contradicted claim

Prediction Didn’t hold up
Angelopoulos: LMArena will launch Data-Driven Debugging within months
“So we're building a project now that we call data-driven debugging D three. It's, you know, it's a little farther out. It'll come in a couple months.”
Anastasios Angelopoulos May 29, 2025 ▶ 1:26:14 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable

Expressed certainty vs assessment result

none yet certainty 1
none yet certainty 2
none yet certainty 3
75% certainty 4
100% certainty 5

weighted support: a fully supported claim counts one, a partly supported claim counts half. Each filled bar is clickable and opens exactly those claims; "none yet" means nothing said at that certainty level has resolved yet

How they sound: speaking style how? →

293 words/min while actually speaking · 7.9 um and uh per 1k words

No argument clarity score for Anastasios Angelopoulos: no usable question→answer exchanges on raw tape (a fair score needs 8+). We do not score a sample that small. Roundtable and news formats yield far fewer direct exchanges than interviews.

Measured by listening to the audio itself: 6,491 words across 1 episode of raw-level tape, transcribed verbatim with every um and uh kept, each one attributed only where the alignment onto our timed stream is unambiguous. These are measurements of speaking style. We do not rank them: across this corpus, fluency and argument quality are nearly uncorrelated (ρ≈0.2), and smooth talking does not signal clear thinking. How it's measured →

Everything Anastasios Angelopoulos said on the a16z Podcast that made the record, most notable first. Filter by type, assessment or year in the ledger →

Opinion
Angelopoulos: Believing chatbot leaderboards are easily gameable is naive
“I have to say, I also just like completely disagree with the foundation of the question. The like implicit assumption is that like chat is easy or that it's even easier than web dev. That's completely false. It's a completely naive perspective that people have…”
Anastasios Angelopoulos May 29, 2025 ▶ 24:57 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Insight
Angelopoulos: Most mission-critical AI queries are subjective, not factual lookups
“In reality, even in such industries, the majority of questions that people ask are subjective. Okay. So the mythology that in hard sciences or in mission critical industries, people just have like cut and dried questions and they just need like a retrieval and…”
Anastasios Angelopoulos May 29, 2025 ▶ 2:50 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Assertion Not checkable as stated
Angelopoulos: Chatbot Arena is immune to model overfitting by design
“Static benchmarks overfit. Why? It's because as Jan said earlier, you're giving the student the same test over and over. You have a model, you test it, you know, you look at whether or not it's improved on a static data set. Then you find another model, you te…”
Anastasios Angelopoulos May 29, 2025 ▶ 21:00 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Assertion Supported
Angelopoulos: LMArena makes style control the default AI evaluation method
“That's why we're making style control default.”
Anastasios Angelopoulos May 29, 2025 ▶ 12:55 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Prediction Not checkable as stated
Angelopoulos: Future AI evaluation will shift to personalized user leaderboards
“Absolutely. Absolutely. It should be personalized just for you. You should understand which models are best for you.”
Anastasios Angelopoulos May 29, 2025 ▶ 26:06 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Opinion
Angelopoulos: Industry AI evaluation platforms face skepticism over bias
“The fact that we come from Berkeley and from a university really speaks to our scientific approach in neutrality. I think if it came from an industrial lab, people would always have questions about, oh, well, these people are they also training a model and wha…”
Anastasios Angelopoulos May 29, 2025 ▶ 39:36 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Insight
Angelopoulos: AI evaluation performance follows a data scaling law
“Because language models are sort of the intermediary that gets you to this evaluation, there's also a scaling law that comes along with it. Which is to say that the more data you get, the bigger you build the platform, the better you can make your evaluations,…”
Anastasios Angelopoulos May 29, 2025 ▶ 56:18 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Insight
Angelopoulos: AI leaderboards can utilize any form of interaction feedback
“Pairwise comparison feedback is not the only kind of feedback that we can use to construct leaderboards. We can construct leaderboards with any form of feedback.”
Anastasios Angelopoulos May 29, 2025 ▶ 1:26:25 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Assertion Supported
Angelopoulos: LMArena router model outperforms all constituent models on Chatbot Arena
“When you train a prompt to leaderboard model, which is like, let's say a seven billion parameter model, and then you use it to route on just questions on the arena and everybody's questions, that model does better than any of the constituent models that were u…”
Anastasios Angelopoulos May 29, 2025 ▶ 1:29:10 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Assertion Supported
Angelopoulos: LMArena prompt router yields double the performance per dollar
“Now, if you trace the performance, the best performance that, you know, any individual model can give you as part of the router as a function of cost. That's like two X worse than the router. In other words, the router is giving you double the bang for your bu…”
Anastasios Angelopoulos May 29, 2025 ▶ 1:30:12 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Insight
Angelopoulos: Top AI researchers avoid companies building purely proprietary technology
“The best people don't want to hole up at a company and develop a bunch of proprietary technology that, you know, is never going to be released.”
Anastasios Angelopoulos May 29, 2025 ▶ 1:36:34 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Insight
Angelopoulos: High refusal rates do not make AI models inherently superior
“It's not necessarily the model that's like most, like refuses the most to answer these like queries that people ask necessarily better. Some people want a model that's more controllable. Some people want a model that's going to say whatever they want. Some peo…”
Anastasios Angelopoulos May 29, 2025 ▶ 1:42:49 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Assertion Open · timeframe May 2026
Angelopoulos: Human evaluators prefer longer AI responses given equal content
“It's true that people vote for longer responses, you know, preferentially over shorter responses, even given the same contents or well-known human bias.”
Anastasios Angelopoulos May 29, 2025 ▶ 12:31 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Assertion Not checkable as stated
Angelopoulos: LMArena measures user preference, not AGI progress
“We don't claim to be an AGI benchmark. We are faithfully representing the preferences of our community.”
Anastasios Angelopoulos May 29, 2025 ▶ 18:08 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Disclosure
Angelopoulos: LMArena is building personalized AI leaderboard tools
“Yeah, and we should be giving you the tools to do that, and we're currently building them.”
Anastasios Angelopoulos May 29, 2025 ▶ 25:56 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Assertion Supported
Angelopoulos: Bradley-Terry models converge for AI evaluation, unlike Elo scores
“Okay, let's move from Elo to Bradley Terry because we're actually performing an estimate here instead of just like You know, and the ELO score moves over time. It doesn't converge, but Rally Terry models converge and how do we then construct confidence interva…”
Anastasios Angelopoulos May 29, 2025 ▶ 38:49 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Insight
Angelopoulos: Reinforcement learning allows AI models to surpass human teachers
“And supervised learning, you can only do as well as the best human that you have. Because what's happening is that you're learning from the teacher. In reinforcement learning, you're learning from the world. You're able to learn things better than the best hum…”
Anastasios Angelopoulos May 29, 2025 ▶ 1:00:48 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Assertion Not checkable as stated
Angelopoulos: Chatbot Arena has 1M+ monthly users and 150M+ conversations
“A lot of people don't know this, but ShopBot Arena is Used by like a million plus monthly users. We get like, you know, tens of thousands of votes on a daily basis. We have like over like, you know, a hundred fifty million conversations that have been had on t…”
Anastasios Angelopoulos May 29, 2025 ▶ 1:13:19 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Prediction Held up
Angelopoulos: LMArena will remain open-source as a commercial company
“We're going to keep publishing papers. We're going to keep releasing open source. We're going to keep releasing open data.”
Anastasios Angelopoulos May 29, 2025 ▶ 1:36:21 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Prediction Not checkable as stated
Angelopoulos: Real-world testing will remain fundamental for evaluating AI agents
“The fundamental is organic, real-world testing with feedback. That's not going to change. I can tell you that that is not going to change.”
Anastasios Angelopoulos May 29, 2025 ▶ 1:44:12 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Disclosure
Angelopoulos: LMArena conducts pre-release model testing for AI developers
“One of the things that we help everybody to do is pre-release testing of their models. Okay. So it's not just that, you know, we work together to evaluate the models are released, but we also try to be their release partners and say, Hey, can we help you guys …”
Anastasios Angelopoulos May 29, 2025 ▶ 4:38 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Prediction Didn’t hold up
Angelopoulos: LMArena will launch Data-Driven Debugging within months
“So we're building a project now that we call data-driven debugging D three. It's, you know, it's a little farther out. It'll come in a couple months.”
Anastasios Angelopoulos May 29, 2025 ▶ 1:26:14 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable

Appearances (1)

EpisodeDateSpeaking time
Beyond Leaderboards: LMArena’s Mission to Make AI Reliable May 29, 2025 28m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.