The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Alex Atallah argument clarity score 4.3/5 from 19 exchanges on raw tape · average scores: directness 4.5 · coherence 4.4 · precision 4 · compression 3.8 record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
19exchanges match
19on raw tape
2redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Can I ask you, when you go back to the founding thesis of the company, what has happened in the ecosystem, in the model landscape that you did not expect to happen?

A Um, okay. Well, one thing that we did not expect was that a, an ecosystem of companies would emerge to host and serve the open weight models. Um, like early on, it wasn't clear that, that that market wasn't going to be a monopoly where like just, you know, the three hyperscalers serve all the open weight models and, uh, and start, and startups don't, you know, there are, they're really far behind. In reality, like, you know, how often do you hear people running, you know, GLM on a hyperscaler? Never. Like they're using the inference providers like fireworks and together and, Um, there's like a big list that we, that we see doing the best job of hosting all the open weight models. And, um, in the early days, We, um, we had, I think we called it provider one and, and provider fallback. We didn't like show which providers were actually doing the hosting because we weren't really a marketplace. We were kind of a, like we were an exploration tool for like finding and discovering new LLMs. And we wanted to get, we wanted to build like a marketplace of model labs, but like the inference provider layer, we weren't sure would actually be a marketplace. And it turned out that those companies were doing a way better job than the hyperscalers, were way faster to host the models and figure out these edge cases to hosting them. And, um, and uptime was just going to be a, a constant problem. …

AI assessment note: “one thing that we did not expect was that a, an ecosystem of companies would emerge”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q We've seen token prices fall. 90% people take in like 18 months. Is the reduction of token prices helpful or hurtful to your business? Because obviously you have a take on spend. If they come down and spend is more efficient. Seemingly it's bad for your business. You have a shrinking pie to take from.

A Well, a lot of people talk about the Jevons paradox, but you know, when When, uh, prices go down by 10 X, the usage increases by more than 10 X. Um, but like, no one has really done a great job modeling it. There are, uh, we do have a lot of spot stories that confirm it. For example, Uh, GBT, 5.6 Luna on open router. Um, open AI cut prices by five X and then in coordination with us by another two X. So in total price, the price of Luna has dropped 10 X on open router over the last two weeks. And, uh, guess how much usage has grown? 13 X. So it's a close to perfect Jevons paradox story where you drop prices, 10 X and usage grows by more than 10 X, um, just a bit more. And, uh, and the, also the, the usage is pretty stable. Like it like grew, you know, it like flattened out, but at, at 13 X and then, you know, it's been kind of like growing at the same rate that it was growing before it hit the, the 13 X multiple. Um, So that's pretty interesting, and it's a pretty, like, uh, low variable. Like, there are, there are a few other confounding variables in the story, and, and it was also done in the middle of Deep Seek launching and having a really, really good price, and GLM having a really good price. Like, now Luna is being used more than GLM on Open Router. GLM used to be, like, one of the top, like, Three, four models by token volume, and now Luna is past it. This is the first t…

AI assessment note: “when prices go down by 10 X, the usage increases by more than 10 X”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q I say, oh, the top five models, when I look at open router are all Chinese. What does that mean? They'll go, oh, well, Harry, no offense to open router, but like, it's not reflective of the market, and most people who use Frontier, it doesn't go through that. Like, they use Frontier APIs, and so it's not counted. To what extent are your rankings reflective of true token usage?

A We try to estimate how they're off, um, by, you know, just surveying people sometimes or looking at like the surveys other people have done. Um, I think we, we have a, we definitely have a bias to People who believe our thesis, which is that the future is multi-model and companies who want multiple models. And there are still companies out there. I basically rarely, very rarely run into them now, but there's still companies out there that are just like, oh yeah, we're an open AI shop. Like we have one, you know, we only do open AI models. And so we're not going to see any of those companies. Um, and I think those companies are primarily focused on Like the, you know, the hyperscalers, OpenAI, Anthropic, and Gemini. Um, so we do probably like under count the, the frontier models. Um, but I think like over time, our thesis becoming more and more common to see in other companies in the moment that like, they're like, oh yeah, like we need to use other models. Um, then our data becomes more representative. And as we scale up, the data becomes more representative in general. Um, so my hope is that like, That we, that the, it just becomes like better and better data over time.

AI assessment note: “so we do probably like under count the, the frontier models.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q What's the craziest thing that you see in your seat on top of everyone's usage that you don't think people talk about enough?

A I mean, a lot of companies are obviously worried about cost management and, uh, and freaking out about the amount of inference they're spending and they don't know how to think about it. It's like a whole new way of, of like doing business and thinking about your, your OPEX. Like the old way of thinking about how your, how much you give your employees, you like give them a salary and you kind of forget about it. Like someone knows what, what everyone's making, but like, It's a static number that, like, gets readjusted on, on a quarterly basis maybe after performance reviews. Really, your, your employees all cost totally dynamic, different amounts now, and, uh, I think a lot of, like, companies are putting it on them to do routing, and I think in the future there's a good chance that it will, like, get pushed downwards to the employee level. Your employees should, like, figure out which tools and models to use that are best for their tasks, and then we should figure out how much you're costing, like, due to the choices that you make as an employee. And, you know, your, your cost as an employee is going to be a dynamic number, and it's going to be, you know, dependent on how much that employee is, like, effectively using, you know, expensive and cheap models to do their job. And, uh, and then I would, I advise companies to kind of like still do their normal management work, like …

AI assessment note: “Really, your, your employees all cost totally dynamic, different amounts now”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q said. And she kind of corrected me that a token is not a token, actually, because one provider can make a token go so much further than another token. It's like, how do you get to the store where you can drive around the whole block, or you can drive straight to the store? Tokens can be made more efficient and go further, and that's the job of the provider.

A Yeah, I, I, um, I agree with that. I think that, you know, in some ways we are providing a service to help people discover providers and, um, and, and like ultimately when, when one provider is making a token go further, we, we spend an enormous amount of time on our router, central router tech, So that, that provider immediately gets more traffic as soon as, as soon as we detect that, like there's a quality improvement or a speed up or a price reduction happening, um, immediately starts getting more traffic. Uh, and the stuff happens like 24, seven, every single, like every, every five minutes, there are big changes for the big models. Um, and so it like actually does make the experience better.

AI assessment note: “when one provider is making a token go further, we spend an enormous amount”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Do you feel a sense of responsibility for that? And what I mean by that is like, you know, you are a routing business, and you could route a company to a Chinese model that, who knows, people are worried about back doors, ICCP involvement. You could be the The deliverer of that to those models. Do you feel a sense of responsibility for that?

A So we, we do feel a responsibility to have safe access for all of these models. Like, um, customer trust is like our, you know, paramount goal. Uh, if, if one of these models is unsafe to use, you know, generally considered unsafe, we'd pull it from the platform. If there's like a way to use it in an unsafe way, I mean, there's a way to use like all the models in an unsafe way. And then we believe in using technology to make it safe. And to like work with the model labs themselves to figure out how they're doing it on their side so that we can be state of the art or better. We spend an enormous amount of time, um, making sure that like that our practices like match what the best things that we're seeing coming out of the labs or are better. Um, and because we're like a very Good. Because we're, we're a way of like exploring all the models and finding them for the first time. Um, we're a good focal point for like deploying safety measures across your whole company. For example, we have prompt injection protection. You can just turn it on and immediately flag prompts that look like prompt injection. Um, that's trying to happen. Um, we have PII redaction. We have Uh, we have like a couple different things that you can automatically just turn on with a click and, and get an added safety layer on top of all of your inference. Um, and we build that so that enterprises feel like they …

AI assessment note: “we do feel a responsibility to have safe access for all of these models”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q And the thing I think is like, What about loyalty? And you have this incredible seat in the ecosystem where you can see everything. Do we see any developer loyalty today with models?

A Honestly, we do see some, we, you know, we, we try to make switching costs close to zero so that when new models come out, um, people can try them out really easily, but we also measure retention and churn from all the models. We share this with this data with model labs too, when they ask for it, um, so they could know like, oh, you know, for my model that just came out, like which models drove traffic to it? And like for those users, like when they leave, which models are they leaving to? Um, and we'll like make this more and more, uh, available to the, to the world, um, soon. And we do notice in the churn data, there are developers who kind of like continuously stick to models, even when there are better models out there, better models for their use cases. Um, I think it's a combination of like a couple, Probably root factors. One is like my app works and I don't want to break it. You know, if the support bot starts saying something weird that I didn't expect, like why, why add more headache? I've already done all this optimization and like, I've already put all these guardrails around it. Um, another is, uh, new, new models are not necessarily going to make your pricing better. In fact, um, in general, what happens is that the current models like price goes down over time. And especially when new advancements in, in, uh, in the labs happen, you know, you'll see like intelli…

AI assessment note: “Honestly, we do see some, we, you know, we, we try to make”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q People thought before that memory would be the retentive mechanism. Well, OpenAI has all of my previous, like, prompts. It knows that I live in London, I do podcasting, and that will make it a better model for me moving forward. Is memory no longer a retentive mechanism?

A Memory is really interesting. I, um, I've always thought it, like, Is a retentive mechanism, and the question is where it lives. Is it going to live with the model? Is it going to live with the inference provider? Is it going to live with the app? Is it going to live with the infrastructure provider, the router? You know, my, my guess is that all of those layers are going to try to own memory in different ways. Um, there, and there are going to be advantages to sticking your memory in each Layer. You know, if you stick it with the app, then the memory has like the most app related context and is model agnostic. If you stick it with the model, the memory might perform the best on personalized benchmarks and perhaps have the best, like ultimate intelligence. And I think the model labs are Work on memory. And, and then the ultimate thing might be like, is there a good combination? Like, can I use memory in the model and memory of the infrastructure layer or the app layer at the same time? Like, is that going to confuse the model? We don't know yet. I do think that like, it's impossible for one layer to capture all valuable memory because the apps own so much important context that the model labs don't have. Um, And they, the model labs in order to get this to work, they'll have to incentivize the apps to, like, give them that context.

AI assessment note: “I've always thought it, like, Is a retentive mechanism, and the question is where”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q A lot of people suggest that that inference provider layer is a commoditizable element or layer that will be removed or see margin reduction competed out over time. What would you say to that theory?

A Right now we're in a massively supply constrained market where, um, and it's likely going to be supply constrained for a while where all the inference providers are, are short, pretty much constantly short. Uh, and, and you're like, okay, so GPUs are, are really, really beneficial. And like, why doesn't Google or Amazon or Azure run around and like buy up all the GPUs and, Take all these inference providers out of business. Well, the people making the GPUs don't want that. Like one of NVIDIA's top priorities is not having customer concentration. They want lots of customers to all have like separate, like allocations of GPUs. Um, they want the, the heterogeneity of the market. They want like competition on the compute layer. Um, and I, and this is good for the ecosystem. Like, like users also want this. This is like, this is, it's good for Nvidia and it's good for end users as well. It like allows these inference providers to kind of like come up with new innovations on like how to serve the models better. Even a single model like Kimi K three, um, uh, like Moonshot just posted a benchmark showing all the inference providers and how well they're serving Kimi K three. Um, and, uh, the, the, they're, they're pretty different numbers for like, for benchmarks that are really static, that are well known. We post this continuously all the time. We always are like benchmarking all the …

AI assessment note: “allows these inference providers to kind of like come up with new innovations”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q How significant was the latest Kimi model, which got so much attention? Was it as significant as everyone thought?

A It's quite good. It's, it's, uh, it's not cyber capable in the way, the same way the frontier models are and long range, long horizon tasks. I think it's still a bit behind, um, the frontier models, but. GLM 5.2 was a really big, big step for open weight models. Kimmy was kind of like moonshot getting up to that step. That's a little bit how I see it. Um, and Kimmy's also a very good writer. Like the voice and tone are both pretty good. Um, Whereas like some of the frontier models, I have like voice degradation that happens when they get better at coding, especially. And I was like, oh my God, like the, I can't, I can't read this output anymore. The output sounds like three of the four arguments you made are right. And one is a turning point, you know, and here's the rub. Like I, it's just, it's sometimes just impossible to read what, what they're saying. Um, Um, and, and this stuff is fixable, but, uh, but Kimi, I think, has always had pretty interesting writing.

AI assessment note: “I think it's still a bit behind, um, the frontier models, but”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q Is that not what APIs did for apps?

A Yes, but, uh, it's much more reliable and, and, uh, deterministic and sort of easy for users to grok with a harness because, um, the harnesses are Unix based. They all have, and, and the models are so well trained on, on Unix, on bash commands. Whereas like, You know, if I'm like telling a harness to go orchestrate an app in the cloud, it's going to be like, oh boy, does this app, like, how do you log into this app? Is it like, do I need your password? Do I need a, uh, do I need to fire up a virtual browser? It's going to be pretty slow. I'll figure it out. Okay. I, I like fired up a browser and like, now I need your password and I'm going to like try to find the input where to put it in. And apparently there's probably an API in this app somewhere. I need to like look up the docs to Figure it out. And okay, now I've got the API, but there's so many like unknown unknowns when you're composing around an app. Very, very, very, very few unknown unknowns when you're composing around a harness. So I just, I think it just gives developers more flexibility, um, and flexibility that they can inspect like API calls. You're just seeing a whole bunch of code flying around the screen, a harness. Oh, I can like jump into the harness and like, look at what's going on and talk in English. About it. So it's much more user friendly.

AI assessment note: “Yes, but, uh, it's much more reliable and, and, uh, deterministic”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q Do you think they will be a serious challenger?

A I do. I think they're, I think they, they're, they have the resources. Um, there's, uh, I think there's some competitive things they can do, uh, around the model that like helps people in ways that the, the model labs are not as interested in doing, like just having like a social network Um, and like a focus on people, uh, you know, it, it's like something for the brand that maybe Grok and like space, like X, SpaceX AI have it too. They do need to find their niche. Like I'm not quite sure. I think people don't quite know what to do with Muse Spark yet. Like when to use it or when to go for it or what it's like, like core advantages. Like they just released a coding harness. Uh, they are, they're trying to be like a generally capable model right now. Um, I expect that in the future, they're going to be like, look, we are way better at this thing. And that's, that's going to be a really important moment for them.

AI assessment note: “I do. I think they're, I think they, they're, they have the resources.”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q I want to create a open American ecosystem. Yeah. More amazing open American models. And I make you head of this program. What would you do to encourage, incentivize the open US ecosystem to compete more vociferously with the Chinese?

A I think I would spend time talking to the current American labs a little bit more to figure out what distilling the Chinese models looks like for them and how effective it is. Um, you can, you can probably get pretty far distilling the Chinese models. Um, the nice thing about the open weight models Uh, and, uh, and the Chinese models that they allow distillation and they like most of them. And that means that you can like take the outputs of these models to do reinforcement learning on top of the model that you're building. And, and, you know, this is just like a very important and common practice in AI that all labs do. Um, so. You know, like it's, it's one, I, I would want to learn a little bit more about like how effective it is. Um, but it's one way to catch up, uh, with the open weight models, um, because they allow it. The other thing is they, uh, the, the American models, and also when you distill, you see the output, so you can like inspect them to make sure that they're aligned. So if there's anything about the, like, you know, the open weight models that you're worried about, um, Not being aligned with like the voice or constitution of the model you're creating. Um, you have a much better shot at catching it when you're doing these RL rollouts. Uh, the, the, the other thing I would try to figure out is Um, is the compute question. Like, compute is just a huge advantag…

AI assessment note: “I think I would spend time talking to the current American labs a little bit more”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q What do you make of every company kind of posturing, haha, we hacked someone? First you had OpenAI, then you had Anthropik, and then you had Zark coming out, I don't want to miss the party, we did too.

A Yeah, well, I think they have, they, they have to talk about it, like, The right thing to do is to reveal when there's been a cyber incident involving your model. Um, covering it up doesn't work. It's not going to work in the long term. And it certainly looks like they're all bragging about it. Um, but really if you were in their position and, you know, something happened with one of the models, um, and you had to make the choice about whether to publish it or not, Like, I think the right thing to do is to publish it regardless of what, like how people are going to spin it. So I don't know. I don't like, I, I doubt that, that they're, you know, they're actually, you know, thinking of the felony bench or.

AI assessment note: “The right thing to do is to reveal when there's been a cyber incident”

Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q immense time that you spend on the routing technology that you have. A lot of people are thinking that we're seeing the commoditization of The routing technology. You're seeing ramp release products like this. I mentioned earlier of, uh, merge company we invested in has released that product. Um, several are releasing kind of routing technology similar or claiming to be similar. Are we seeing the commoditization of this layer?

A I think a lot, yeah, a lot of companies are making routers because it's fashionable. Um, I think they're, you know, they're seeing growth happen here and, uh, or they're making gateways at least. Um, first, I think there's two, there's two issues with that. First, it immediately puts you in the mindset of copying instead of like, you know, winning something. Um, you're sort of, you're playing, um, To play or you're playing to exist rather than playing to win. Um, and, uh, and maybe you're just trying to like play to serve your, your existing customer base. Uh, and you, you want to see some AI growth happen. Um, I think, you know, immediately kind of like puts that gateway, like many, many months behind, um, the companies that are fully focused on it. Like I am a hundred percent focused on building the best router and gateway. And, and LLM marketplace. Um, and it shows in our product and, um, and you know, the, the benchmarks that we create internally and how we see ourselves compared to the competition. Um, this is not a side quest for us. Like it, it may be for some other companies. Um, the other problem is that it, uh, it reduces the leverage of all of your users. So, um, Like, I really deeply believe in giving users and developers more leverage. Like, fundamentally giving them access to more models is about giving them more leverage over all the innovations that happen in AI…

AI assessment note: “puts that gateway, like many, many months behind, um, the companies that are fully focused”

Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q Can I ask you, Alex Kopp said on CNBC in his rather wonderfully energetic way that companies are terrified of working with frontier model providers. Do you think they are?

A I haven't seen what he has, what he talked about there when I talked to our customers. Um, but there was definitely like a little, there was some skittishness that the, Uh, particularly when, when Claude design came out, uh, around Figma. And that part, um, I did see, and I do think that there are like real concerns for a company that, um, is kind of building like a, you know, thin, you know, like Figma is very different, but if like a, a startup is only building, um, a, Like go to market wrapper around intelligence. Like, Hey, we are, you know, we're a company that kind of like brings AI to this market and does so by like doing the right integrations and like customizing the system prompt. Um, you're gonna be fine if the model labs don't care about that market, which There'll be many markets like that, but the model labs have several incentives to go after you eventually. One is, uh, getting multiple teams within companies they do care about to be dependent on them. So this is my theory behind why like Claude design was strategic. While it's not like a massive amount of revenue for Anthropic, like not probably not a significant amount of revenue. Um, it does get the design team to really care about Anthropic models. And so the companies that like they want, like they now have another team that really wants to stick to Anthropic. So those team, that team strategy, Uh, makes thi…

AI assessment note: “there was definitely like a little, there was some skittishness”

Answered raw tape D 4 · C 4 · P 4 · Cm 3 3.85

Q What will be the main revenue line of open router in three years time?

A I mean, I think it's going to depend on, on the economy in so many ways. Like if the overall AI market keeps growing the way it's been growing over the next four years, you know, you know, with like 10 to 15 X every year, um, Or potentially more. It's a lot of growth. Um, you know, I think under that, under that world, um, I would expect people to continue to underestimate how much inference they're going to need. And thus our, our revenue is going to be dominated by, you know, the same things that dominated today, that dominated today, which is like, you know, like us helping people with, With an unplanned inference capacity, um, both enterprises and startups. And, uh, like that, that's what open router is best at. Like when you, when you need to try models that you weren't expecting, you need to try when you're like, um, using more inference than you thought you were going to use on particular models. Like we make, we make sure that that is not going to be an issue for your company, um, by providing the best failover and best uptime. And this is really, really a good thing to do when the market is like continuously underestimating its inference needs and growing at this rate. Um, if this, uh, growth rate continues over the next four years, um, I mean, it's going to be a wild amount of growth, like the economy and the economy has some limits to it. I can see like Major SMB, yo…

AI assessment note: “our revenue is going to be dominated by... us helping people with, With an unplanned inference capacity”

Redirected raw tape D 2 · C 4 · P 4 · Cm 3 3.25

Q Speaking of the apps and the model set, claw code, cursor, bundle, model, and harness, is the router absorbed into the agent framework before it ever has the chance to be independent when you have the agent and the harness together?

A The harnesses are pretty interesting because in our early days, one of our early bets was that Most apps were underestimating the desire for users to choose the model. Like most apps in the very early days in like, 20, 23 and 20, 24. It wasn't even clear which model was being used under the hood. They were like, ah, people are not going to care about that. They just want AI. Um, and our, and one of our like strong convictions then was that no, like people are going to want to like use particular models. They're going to care about who they're talking to. It's like, you know, I, I want to know which employees I'm talking to when I'm trying to solve a problem and models will be kind of like that. Um, and that has played out, you know, like in notion, you can like choose the model that you, you talk to. Um, even though you would think an app like that might want to like obscure it completely. Um, similar, a similar thing happened with harnesses where particularly with developers, um, they started to build an affinity to different harnesses. And, uh, and that's cause it's like, it's a user experience. So I think that is my favorite argument for why harnesses are going to stick around. Um, not that like, they're being bundled with the models, because in fact, like, As models get better, they get more resourceful and the, the junk that gets thrown in the system prompt just becomes a …

AI assessment note: “The harnesses are pretty interesting because in our early days, one of our early bets”

Redirected raw tape D 2 · C 4 · P 3 · Cm 3 3.00

Q You can only invest in one inference provider. Which one do you invest in?

A I probably have to stay, you know, stay neutral on this. I, I do, I really like the, the, the, you know, inference providers that are doing, um, that are doing like custom hardware and, uh, and very, very like low level optimizations. Um, I like providers that are also trying to figure out how to make, um, customization easier. So like today you fine tune models and, um, and you create this like new, like fully independent model from the, from the base model. Um, many inference providers are kind of like, you know, creating these Laura's or some, some call them like cartridges that are much more portable. Potentially, between models, and we might see a future where, like, when you do a fine tune, and you want to, like, change the base model layer, it only costs, like, maybe a few hundred dollars, maybe a few dozen dollars to change it.

AI assessment note: “I probably have to stay, you know, stay neutral on this.”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.