The Exchanges, every show

Every argument clarity score on this site is built from rows on this page, here across all 44 shows. Each question and answer was assessed with names hidden, the hosts' own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

shows every show 44 of 44
every show
Anastasios Angelopoulos no published score: a fair score needs 8 or more exchanges on raw tape on one show record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score rests on one show's raw tape, the show with the most assessed exchanges, and shrinks small samples toward that show's cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
21exchanges match on 44 shows
21on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q Well, that was incredibly succinct. Thank you. Normally people take about four hours after I ask for a succinct description. When you look at the models that you have on ARENA, the sheer number of them, Bluntly, I just am faced with the one question. Holy shit. Is this like the true commoditization of models? Are they just a complete utility layer at this point?

A Well, I think that there's, ah, the, the big question around this has started to rise because of open source models. So I think if you were to only look at the closed source models, you would say there's acceleration, but it hasn't quite commoditized yet because that layer is still owned by a pretty small group of companies. It would be an oligopoly if we only had the closed source models. But what seems to be happening is that the open source models, especially from China, Have really rapidly improved. And for the first time ever, we saw a couple of weeks ago that Kimi K three actually beat the best closed source American models, uh, on a, you know, pretty important subset of tasks, for example, front end, front end coding, like web development, which a huge fraction of developers are web developers.

AI assessment note: “if you were to only look at the closed source models... hasn't quite commoditized yet”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Should they, should they be restricted in terms of access to US markets? Because what's funny is the US is like, oh, should we restrict access? And the Chinese are also going, oh, should we turn them off too?

A Totally. And by the way, it's worth noting that China has already restricted the use of American models within China, right? So if you look at the two by two matrix of US China restrict, not restrict, you know, like export import stuff, um, They, they have already restricted the use of US models within China. It's only Chinese models that can be used in China, which affects all American companies. And so then there's the pro cons of all sides of the following regulation. If China restricts the use of Chinese models in the US, what are they giving up on? Revenue and global mindshare and dominance. That doesn't seem like a good trade to me. And then what are they getting? In return. In return, they're getting that the US doesn't get to benefit from Chinese open source models, which of course would cripple American businesses in the sense that it wouldn't allow them to build on the best open source intelligence. At the same time, it would make open AI and Anthropics stronger. Right? So that is kind of the trade off on the Chinese side. I don't really see them banning the use of Chinese models in the U.S. I don't think it makes sense for them. And then on the other side, should the U.S. ban Chinese models within? I think that there's also trade offs. So on the pro side of banning, there could be back doors in these models that are dangerous. And it could, by banning, we could allow…

AI assessment note: “should the U.S. ban Chinese models within? I think that there's also trade offs.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Okay, so what, so why is it interesting?

A So I have a thesis on hypergrowth. There's two types of hypergrowth markets that we see today. Market A is what I call scaling complements. And these are goods that are complementary goods to the scaling of AI models. And I mean that in the economic sense, a complementary good is good A and B are, the good A is a complement to good B if the demand for good B drives demand for good A. So if I have a car, gas is a complementary good to cars. The more cars are sold, the more gas is sold. And so data is one of these scaling complements. Because the bigger models scale, the more data you need, and that's a scaling law question. And so the more models you get, and the bigger that they're getting, the more they're proliferating, the more businesses are training their own models, the more data you are going to need. And it's a very fundamental need. People forget this. They think about data as a commodity. It's really not. It's actually less so of a commodity than even GPUs. Because in order for data to become, uh, irrelevant, Humans need to become irrelevant, and that means that we've achieved AGI. So data is a very durable need, and companies are spending on it, usually within Frontier Labs, at about 10 to 20% about the amount that they're spending on GPUs. And so if you believe in the GPU market accelerating, if you believe in the scaling of models, if you believe this is going to b…

AI assessment note: “data is one of these scaling complements. Because the bigger models scale, the more data”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Dude, can I ask how big a moment was that? Because it's like a, I'm going to butcher it, but you know, I'm a podcaster so I can get away with it. You're a, No, it was a pretty big moment.

A It was a pretty big moment. It was a pretty big moment. And the reason I'd say it was a big moment is because it violates A narrative that has been persistent in the United States, which is that the Chinese are just distilling American models, and that's the only way that they're able to, you know, keep up. When really, what happened is that Kimi actually beat all American models, including Fable, in some subset of tasks. That doesn't mean that they're not distilling. They may still be using distillation as a sub-step in their training procedure, But it does mean that distillation is only part of the story and that there's something that those labs are doing above and beyond distillation that's bringing the performance up above what the American labs are, are currently doing. And so that narrative violation has been hugely important to the way that people view the ecosystem, both from the scientific dominance of Americans and the American sort of hegemony Uh, which of course Americans love hegemony, uh, to, um, the, uh, economics of the whole thing. And to your point, are these things, are these models a commodity or not?

AI assessment note: “It was a pretty big moment. And the reason I'd say it was a big moment”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q When we see, you know, when we look at your open readers of the world, you know, the top five models are all open source Chinese models. When we see the proliferation of Chinese models today, does that cannibalize the closed frontier model business meaningfully?

A Well, I think that you need to think about the, um, the incentives and economics behind it. So first thing I'll say is that the open router metrics are not truly reflective of reality, and that's because the business route, the business model of open router is to charge, uh, like a fraction, uh, uh, like a fee on top of every token. And so what happens is that people don't use open router for proprietary models. People are using open router primarily for open source models where they need the failover and all the Value-added services that OpenRouter provides. If you look at the whole space of all inference, most of it is still being consumed on first party APIs and on proprietary models. That's why anthropic revenue has been just a total hockey stick. It's not, you know, it's not like they're being completely cannibalized right now by Chinese open source models. These models are still only a small fraction of the total inference spend in the world. That said, think about what's happening in the future. Enterprises are going to want to own their own intelligence. They're going to want so-called AI sovereignty, which is a fancy word for meaning that you own your whole supply chain of AI. That means you can take an open source model and you can fine tune it on your own company's data and own your stack end to end, basically outside of the compute hosting. And so then you should be…

AI assessment note: “It's not, you know, it's not like they're being completely cannibalized right now”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Why have we not so far? I, I really hope so too, by the way. I completely agree. We'd love to see that. But why haven't we? Why has the US open community lagged behind so meaningfully?

A Well, frankly, I think it's a business model question. You know, I think that people have not really figured out up until this point what the business model is for open source. And now I think people are wisening up to it. There's a few different ways of going about it. One way of doing it is to say, I'm going to do a rev share. I'm going to take this open source model. I'm going to allow inference providers like a fireworks or together or whatever to deploy, um, this model. And then if they get to over X dollars in revenue, I'm going to ask to do a revenue share. And that is one way of building a sustainable company off of open source. You basically share in the compute revenue. Another way of doing it, which is I think the more mystical thinking machines type of strategy is to take the open source model and then use it as a lead generation tool for companies to build on top of that and then come to you and say, can you help us fine tune? Can you help us with our AI strategy? And then you do that for deployed engineer. And that is actually a huge market because if you think about it, one of the biggest markets over the next 10 years is going to be AI modernization. Going into every business in the world and then helping them retool in the face of AI. Take advantage of their data, restructure their data, you know, figure out how to use these models, integrating them into workfl…

AI assessment note: “frankly, I think it's a business model question.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q fireworks and she was like, specialized intelligence will be the future. Companies will have their own fine tuned specialized models with their own company data and the performance will be better. And that is what will happen. Do you think that's right or is that actually just a small subset of very advanced Silicon Valley companies and the known yogurts and every normal company would just use Frontier or whatever?

A Well, I'll say it like this. I think the business incentives make this inevitable. And the reason is because businesses are going to need a way of keeping a moat in the age of AI. Software is no longer really a moat because it can be produced instantaneously, right? Let's project out five years. That's what's going to happen. And so what moats exist? Network effects exist and data moats exist. And if you can take your data mode and turn it into a self improving product, that is a way for businesses to remain sustainable in the age of AI. Let's say I'm a business like a Coca-Cola. I'm a Cisco. I have a lot, a lot of users. I might not be necessarily at the frontier of the AI technology of the world, but I do have this massive corpus of data that I can use in order to, you know, beat my competition. So what should I do? I should be trying to take advantage of my data. As much as I possibly can to accelerate my business and stave off competitors.

AI assessment note: “I think the business incentives make this inevitable.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q it was, it was good. Um, I totally agree with you. Um, okay. Totally get that with all of these models. The question also becomes, huh, what model should I use? We spoke about OpenRouter earlier, and it seems like since the announcement that they were getting bored, everyone just has their own routing product. Is there a value in the model routing layer, and how should I analyze that?

A Yeah, I absolutely think there's value in the model routing there. That's why lots of companies are doing it. And, you know, we'll see which ones end up standing the test of time and which ones are actually a priority for the companies. I think there's an element of hype cycle right now around routing that needs to be kind of like purged before we see, uh, who ends up actually building a great router. But routing is a very difficult technical problem. That's the first thing to, to like realize. Because in order to route, you need to be able to take a query And then you need to understand the nature of the query, how difficult the query is within its domain, um, which, which is hard to tell. And then you need to also understand, based on data, all the performances of the different models that are in the surf set, and also be able to quickly onboard new models that are being released, as we said, every week. So that technical challenge, imagine if every enterprise in the world was trying to build this themselves. They wouldn't be able to do that.

AI assessment note: “I absolutely think there's value in the model routing there.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q What did you do with Arena that, with the benefit of hindsight, you wish you hadn't done?

A Oh man, I had so many mistakes. I mean, at the beginning, I had no idea what I was doing, and I, you know, my co-founder Jan probably knew and could see behind the corners, but I was probably too stubborn to listen to him. So, first of all, I've learned to listen to Jan more. But second, it's, you know, I, so many, like, experiments at the beginning that I just shouldn't have wasted time with. I think the, the degree of focus that you need to run a company It's just so extreme. You really need to do one, maybe two things extraordinarily well, and focus very, very deeply on them. Pick the right ones and focus on what's working, not on expanding into things that are not working. And that is a, that is a great lesson for me.

AI assessment note: “so many, like, experiments at the beginning that I just shouldn't have wasted time”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Do you think that severely impacts their ability? Again, I'm naive. That severely impacts their ability, or we're actually just fostering an ecosystem where they're going to learn to build it really fast because they don't have access to it?

A Well, I think it may be hindering them now, but I think it's a good question as to what's going to happen in the future, because they are really good at building hardware, and the, the, the downside of export control is that, uh, it can incentivize them to build their own ecosystem, and then what do we do? You know, so the hope is that we keep Nvidia ahead of the game so that we can retain the advantage that we have and the TSMC's to the world and our whole, that ecosystem is absolutely a national security necessity. So we should have the government, you know, really protecting it and growing it as well as new companies that are innovating, you know, etch just came out as an example, um, within the United States, uh, to continue to build on our lead there.

AI assessment note: “Well, I think it may be hindering them now, but I think it's”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q A bit of a sticky situation then, aren't you?

A Yeah, totally. I mean, you know, I was just talking with a big, you know, Fortune 50, uh, enterprise, uh, yesterday, and they were, and I was telling them about, you know, products that we have for them and so on and so forth, and they said, okay, wait, is anything in your stack built off of Quen? And I said, you know, yeah, we use Quinn for X, Y, Z. And they're like, is that flexible? Can you, like, stop doing that and use an American model instead? And I was like, oh, interesting. You know, I totally understand where you're coming from. Uh, yes, we can do that. Um, but also, I'm gonna talk to Harry about this tomorrow.

AI assessment note: “Yeah, totally. I mean, you know, I was just talking with a big”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Why would that exert downward pricing pressure? Just because everyone will be like, you can't have that high margin to your price gouging?

A Yeah. People are going to be like, well, I know that you can do a better discount. Like in negotiating leverage, I'm like, okay, like a standard negotiation with a private company goes like this. I'm charging X. And then the other side says, no, you should, it should be one third X. And they're like, I'm so sorry. Like, I can't run a business that way. I'm just going to go home hungry. I need to make my bread too. I hope you understand. Like, I'm not trying to price gouge you. And then the other side's like, okay, two thirds X. And then you're like three quarters X. And they're like, make a deal. But imagine that the other side is full information about the fact that you're charging twice as much as you need to. Then it becomes easier to negotiate.

AI assessment note: “People are going to be like, well, I know that you can do a better discount.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Totally. Does that have to be a neutral, non-company, non-government body that does that regulatory role?

A I think if it's not a company, it's going to be tough. I, I understand the need for something neutral, Um, but, you know, you want, you want to let the incentive system work itself out. So I would say that, like, we, we should create strong, uh, safety incentives for American businesses, and then regulate businesses based on the outcomes Basically, for example, if like open AI is like letting their AI break into hugging face or whatever, they should get like huge fines and huge scrutiny and all that stuff, as opposed to like, uh, having a government process that's in charge of ensuring that this doesn't happen again, which they won't be able to do that. They're not technically capable.

AI assessment note: “I think if it's not a company, it's going to be tough.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q we talked about Dumb as Rocks and doing this show. I'm also an ambassador for my sins. And I meet so many of these people leaving OpenAI, Anthropic, you name it, um, and they all kind of seem the same, if I'm totally honest. Smart people out of great company. How, what will determine the Neolab spin-outs that succeed versus Flame Out with a huge amount of cash going in?

A Yeah, I think that the Neolab thing is really tough. So just to, so that we're on the same page with the audience, like, there's at least 75 Neolabs. And For sure, like two thirds of those are going to be worth nothing or like they're going to be bought out for parts, right? That's going to be like an aqua hire. And so what is going to determine the winners versus the losers in that game? And I think it's all about being very aggressive towards a great strategy and business model. Because what's happened, and you know this better than I as an investor, is that the markets have become very, uh, P&L driven. It's like not enough just to like create a model and then have a party about it. Hey, we created an AI. That is like old fucking news. Today, it's about Not just going to create a model, but do I have a sustainable business model around that? And can I generate hyper growth in revenue? And if you're not able to do that, you're not even going to be able to raise your next round. People are raising multi-billion dollar rounds on top of just the names that are in the Neolab with zero proof that there's any revenue generating model behind that. And so then the question you have to ask is, let's say, I'm one of those people that's at, say, a ten billion dollar Neolab valuation. What do, what do I have to believe in order to 10 X my money? And the thing that you really need to belie…

AI assessment note: “all about being very aggressive towards a great strategy and business model.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q expand that, okay, if we think Anthropic and OpenAI can be three to five trillion dollar companies, let's just put that there. How big does that mean the data providers can be? Like, you know, McCall's reportedly raising now at 20. Does that mean that these providers will be worth a hundred billion dollars? That wouldn't be egregious, would it, to say it's three percent of the market cap of-

A Yeah, I think it could, I think it could easily be a hundred. I think these companies will easily be worth hundreds of billions of dollars, and I think they could even be worth more. The data is really the hardest part of, of model training. Because you need to source it. It's so dirty. Nobody wants to do that shit. Nobody wants to hire all these people to generate data and then, you know, turn that into basically data plus GPUs equals model. And then the algorithms have become somewhat of a commodity because people know how to use the transformer. That's why, as you said, all the people that are coming out of the frontier labs look the same.

AI assessment note: “I think these companies will easily be worth hundreds of billions of dollars”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Isn't value entirely subjective? Like, for one it's speed and for others it's accuracy. Do you know what I mean?

A Right, absolutely. So you can try to decompose it. I, I think about it as three, a three, um, three-pronged, uh, value proposition. There's performance. And then there's cost and latency. Cost and latency are easier to define, but performance is the tough one because the definition of performance depends on the business, depends on the use case. So at Arena, we built this pretty sophisticated pipeline for extracting organic performance measurements from agentic traces. And that's exactly where I would say that the, the value lies in helping businesses take advantage of their own data instead of having to purchase data. In order to say which AI works best for them, um, and even help them train their own.

AI assessment note: “Right, absolutely. So you can try to decompose it.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Was there a moment for you where you were, uh, I'm sure you were debating it yourself. You had other opportunities. What was the deciding factor for you?

A It became clear that the only way to scale what we were building was to build a company out of it. That the world really needed something like Arena. Arena being really a place to sort of measure, understand, and, and advance the frontier AI capabilities in, on real world users, on real world usage. Based on organic feedback and that in order to achieve the scale and, you know, distribution necessary and the quality of course of the platform necessary to do this effectively, we would need to start a company out of it. You know, we considered other options. Are we going to keep doing this as an academic project? Are we going to do it as a nonprofit? Blah, blah, blah. But ultimately under those constructs, we didn't feel like we'd have the resources necessary to accomplish our mission.

AI assessment note: “It became clear that the only way to scale what we were building was to build a company”

Answered raw tape D 4 · C 5 · P 4 · Cm 4 4.30

Q Really smart of them. But like, yeah, your investors become employees in many respects. And so I think they'll lobby incredibly efficiently. The question I have for you is, When we look at Jensen's letter that he did on X, how did you read that? Was that like a incredibly smart realization that he had to do it and it was in his favor? How did you think about it?

A We really believe in the importance of open source to American businesses. Um, and in particular, we believe in the idea of not crippling American businesses by banning open source, but also incentivizing American companies To develop open source models. There's a world where AI is closed source. There's a world where businesses get less choice, higher costs, less competition, you know, more risk. And we don't really want that as a, as an open ecosystem. Of course, Jensen is in some sense self-serving with this letter, because the more open source models are developed, the more companies are going to be training on GPUs. They're going to be fine tuning on their own data. And it's just more and more spend. It decreases revenue concentration of Nvidia. I mean, that business is doing great. They don't need help, but you know, I think it's like, there's, there's a lot of reasons why he should be pro that, um, as should we, but nonetheless, I think, um, it is actually a, a, a patriotic mission.

AI assessment note: “Of course, Jensen is in some sense self-serving with this letter”

Partly raw tape D 3 · C 5 · P 4 · Cm 4 4.00

Q Expert arena. So basically like w what is in the critical path for you, let's say for next year and what, what have you decided you will never do?

A So let me first Talk about things that I'll never do. The platform, integrity comes first to the platform. The, basically the public leaderboard that we show on Ellen Marina, I think of as a charity. It's a loss leader for us. We don't really make money on the public leaderboard. You can't pay to get on the public leaderboard. It's not like a Gartner in that sense. It's not like any of these, like, uh, you know, pay to play systems, never going to be like that. Models are going to be listed on the leaderboard, whether or not the providers pay. And whether or not they're getting a good score. They can't pay to take it off either. And so what that means, that that's very important. And so what that means is that the leaderboard has a certain integrity that will never be compromised, of course.

AI assessment note: “So let me first Talk about things that I'll never do.”

Answered raw tape D 4 · C 4 · P 4 · Cm 3 3.85

Q So that, that's on, like, the data side. Um, can I ask you, when we think about, like, on the agent side, Anjani said I had to ask you, how does your business change as we think about the transition to full trust with agents?

A Yeah, so agents, uh, is the number one priority for Arena and has been all year. People don't know this, but Arena's one of the largest consumer AI apps in the world. We're bigger than, like, XAI. We're bigger than, like, Hugging Face, and Manus, and GenSpark, where it's so massive, like if you, like outside in, it's like 30 plus million, uh, monthly visitors are on Arena. It's, it's, and because, and most of them are knowledge workers and prosumers, people that we call unhirable experts, people that are coming to Arena to do their real daily tasks, and in doing so, they are giving feedback that allows us to build the evaluations that we share with the world, and so it's this organic flywheel for agentic, uh, evaluations based on real data.

AI assessment note: “it's this organic flywheel for agentic, uh, evaluations based on real data.”

Partly raw tape D 3 · C 4 · P 4 · Cm 4 3.70

Q They're awesome. Love, Gavin and team. Totally agree. Can I ask you, just in terms of the export control, do you think it's right that we have the export control on chips?

A I think there's national security questions around these chips. I do think that there, it is a real debate, though, as to which way you want to go about it. Do you want to addict the world to American hardware, which would be a case, the case against export control? Do you want everyone in the world using Nvidia, and therefore That value, basically money into America and then crush competition in China? That would be world A. And then world B would be, is it worth it to cut that off for the short term or medium term impact of us being ahead? And maybe we just like continue to stay ahead and we starve them of the resources that they need in order to build. The regulatory ecosystem around the open source models also is moving in this direction, right? China

AI assessment note: “it is a real debate, though, as to which way you want to go”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.