Everything Anastasios Angelopoulos said on any show that made the record, most notable first. Each card names its show and opens the statement there.
Angelopoulos: Chatbot Arena is immune to model overfitting by design
“Static benchmarks overfit. Why? It's because as Jan said earlier, you're giving the student the same test over and over. You have a model, you test it, you know, you look at whether or not it's improved on a static data set. Then you find another model, you te…”
Angelopoulos: Static benchmarks are intrinsically unable to evaluate generative models
“Static benchmarks are intrinsically, to some extent, unable to measure generative model performance. And the reason is because you cannot Pre-annotate all the outputs of a generative model. You change the model. It's like the distribution of your data is chang…”
Angelopoulos: Forward-deployed engineering plus open-source models will outlast third-party APIs
“I do think the combination of FDE plus Open American model May be a more sustainable model for the future of American or even Western businesses, because it, because they might not want to be building on top of external third party services. They might want to…”
Angelopoulos: Chip export controls risk incentivizing China to build an independent hardware ecosystem
“The downside of export control is that it can incentivize them to build their own ecosystem, and then what do we do?”
Angelopoulos: China Unlikely to Ban Chinese AI Models in the US
“I don't really see them banning the use of Chinese models in the U.S. I don't think it makes sense for them.”
Angelopoulos: Enterprises fear relying on frontier AI labs and Chinese open-source
“It's not only true that they're terrified of working with the frontier labs, but they're also terrified of working with the Chinese open source.”
Angelopoulos: Arena considering mandatory in-person onboarding to verify real humans
“We're gonna change our whole hiring process because of this kind of stuff. It absolutely worries me. Well, at first you need to verify that person is real. So all of our onboarding, we're considering at least making all of our onboarding in person because of t…”
Angelopoulos: Frontier AI labs spend 10% to 20% of GPU compute budgets on data
“Companies are spending on it, usually within Frontier Labs, at about 10 to 20% about the amount that they're spending on GPUs.”
Angelopoulos: Data is the hardest part of model training as algorithms commoditize
“The data is really the hardest part of model training. Because you need to source it. It's so dirty. Nobody wants to do that shit. Nobody wants to hire all these people to generate data and then, you know, turn that into basically data plus GPUs equals model. …”
Angelopoulos: Arena has passed a $100M annualized revenue run rate
“So we're past a hundred million in annualized revenue run rate, and that's based on like Q, Q two times four.”
Angelopoulos: The AI 'SaaS-pocalypse' is overstated due to data moats
“I think that the SaaS-pocalypse has been a little bit overstated overall. Because people don't understand always the dynamics of those businesses and how tough it is to replicate what they've built just also from a network perspective and a data perspective.”
Angelopoulos: Legal AI startups need not fear Anthropic due to differing priorities
“It's like priority number 12 for Anthropic is probably not high enough for Harvey and LaGuardia to be too scared.”
Angelopoulos: AI Will Eradicate Diseases Like Open Problems in Math
“I think that the, like, level of just human flourishing that's going to happen as we start to one by one eradicate diseases the same way that we're currently eradicating open problems in math is going to be incredible.”
Angelopoulos: Value in AI Biology Will Accrue to Data Layer
“That's exactly one of the areas where the data layer, where you can clearly see that the data layer is where value is going to accrue. Because the GPUs Are the same GPUs in both cases. The problem is that, that data infrastructure, the flywheel, the data colle…”
Arena's organic user prompts provide realism that Artificial Analysis lacks
“They have arenas, but the arenas are not based on organic usage. Like the thing that distinguishes our platform versus theirs is that the users are actually inputting their own use case. They're actually asking their own question. And that gives a level of rea…”
Arena sampled open-source models at 60/40, debunking Leaderboard Illusion paper
“But, you know, there, for example said that we were, that we only sampled, like, nine percent open source models and, like, you know, 60%, like, closed source models, and this created a gap between open and closed source. But in reality, we're actually really …”
Arena's public leaderboard will never adopt Gartner-style pay-to-play models
“You can't pay to get on the public leaderboard. It's not like a Gartner in that sense. It's not like any of these, like you know, pay to play systems, never going to be like that. Models are going to be listed on the leaderboard, whether or not the providers p…”
Angelopoulos rejects claims that Cognition's Devin is dead
“Devin's not gone. Devin's everywhere.”
Angelopoulos: LMArena makes style control the default AI evaluation method
“That's why we're making style control default.”
Angelopoulos: Future AI evaluation will shift to personalized user leaderboards
“Absolutely. Absolutely. It should be personalized just for you. You should understand which models are best for you.”
Angelopoulos: Industry AI evaluation platforms face skepticism over bias
“The fact that we come from Berkeley and from a university really speaks to our scientific approach in neutrality. I think if it came from an industrial lab, people would always have questions about, oh, well, these people are they also training a model and wha…”
Angelopoulos: AI evaluation performance follows a data scaling law
“Because language models are sort of the intermediary that gets you to this evaluation, there's also a scaling law that comes along with it. Which is to say that the more data you get, the bigger you build the platform, the better you can make your evaluations,…”
Angelopoulos: AI leaderboards can utilize any form of interaction feedback
“Pairwise comparison feedback is not the only kind of feedback that we can use to construct leaderboards. We can construct leaderboards with any form of feedback.”
Angelopoulos: LMArena router model outperforms all constituent models on Chatbot Arena
“When you train a prompt to leaderboard model, which is like, let's say a seven billion parameter model, and then you use it to route on just questions on the arena and everybody's questions, that model does better than any of the constituent models that were u…”