Assertion certainty 4/5 debate potential 3/5

Arena's anonymous Nano Banana test moved Google's stock and product roadmap

Anastasios Angelopoulos · [State of Evals] LMArena's $1.7B Vision — Anastasios Angelopoulos, LMArena · Dec 31, 2025 · at 13:20

Anastasios Angelopoulos, co-founder of LMSYS Chatbot Arena, describes the market impact of pre-release anonymous model testing on the platform.

0:00 / 0:11exact quote · 11.6s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“I mean, that moment alone changed Google's like roadmap. Market share. Seriously. I mean, Google stock, billions of dollars are moving because of Nano.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Anastasios Angelopoulos

Insight
Angelopoulos: Static benchmarks are intrinsically unable to evaluate generative models
“Static benchmarks are intrinsically, to some extent, unable to measure generative model performance. And the reason is because you cannot Pre-annotate all the outputs of a generative model. You change the model. It's like the distribution of your data is chang…”
Anastasios Angelopoulos Nov 1, 2024 ▶ 6:40 In the Arena: How LMSys changed LLM Benchmarking Forever
Opinion
Arena's organic user prompts provide realism that Artificial Analysis lacks
“They have arenas, but the arenas are not based on organic usage. Like the thing that distinguishes our platform versus theirs is that the users are actually inputting their own use case. They're actually asking their own question. And that gives a level of rea…”
Anastasios Angelopoulos Dec 31, 2025 ▶ 7:24 [State of Evals] LMArena's $1.7B Vision — Anastasios Angelopoulos, LMArena
Assertion Supported
Arena sampled open-source models at 60/40, debunking Leaderboard Illusion paper
“But, you know, there, for example said that we were, that we only sampled, like, nine percent open source models and, like, you know, 60%, like, closed source models, and this created a gap between open and closed source. But in reality, we're actually really …”
Anastasios Angelopoulos Dec 31, 2025 ▶ 11:51 [State of Evals] LMArena's $1.7B Vision — Anastasios Angelopoulos, LMArena
Disclosure
Arena's public leaderboard will never adopt Gartner-style pay-to-play models
“You can't pay to get on the public leaderboard. It's not like a Gartner in that sense. It's not like any of these, like you know, pay to play systems, never going to be like that. Models are going to be listed on the leaderboard, whether or not the providers p…”
Anastasios Angelopoulos Dec 31, 2025 ▶ 17:26 [State of Evals] LMArena's $1.7B Vision — Anastasios Angelopoulos, LMArena
Opinion
Angelopoulos rejects claims that Cognition's Devin is dead
“Devin's not gone. Devin's everywhere.”
Anastasios Angelopoulos Dec 31, 2025 ▶ 23:20 [State of Evals] LMArena's $1.7B Vision — Anastasios Angelopoulos, LMArena
Assertion Supported
Angelopoulos: OpenAI o1 crushed Chatbot Arena, proving the benchmark isn't saturated
“So there's this model and it crushed the benchmark. You know, it's just like really like a big gap. And what that's telling us is that it's not saturated yet. And so it's still measuring some signal that was encouraging point.”
Anastasios Angelopoulos Nov 1, 2024 ▶ 27:20 In the Arena: How LMSys changed LLM Benchmarking Forever
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.