Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q And you have this technique called dynamic multiplexing, which is basically instead of having a one-to-one relationship, you have one GPU for multiple. Clients. And I saw one of your customers, they went from. 30 clients to just one single GPU and they cut costs by 97%. What were some of those learning, seeing hardware usage inefficiencies and how that then played into what you're building now?
A Yeah, I think it basically showed that there was probably a gap with even very sophisticated teams making good use of the hardware is just not an easy problem. I think that was the main, I, it's not that these teams were like not good at what they were doing. It's just that they were trying to solve a completely separate problem. They had a model that was trained in-house and their goal was to just run it. And it, that should be an easy, easy thing to do, but surprisingly still, it's not that easy. And that problem compounds in complexity with the fact that there are more accelerators now in the cloud. There's like TPUs, Inferentia, and there's a lot of decisions that users need to make, even in terms of GPU types. And I guess sort of what we had was we had internal expertise on what the right way to run the workload was. And we were basically able to build infrastructure to make it so that companies could do that without thinking.
AI assessment note: “we were basically able to build infrastructure to make it so that companies could do that”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q And you said, look, what we care about is building things on top of code completion. How did you decide to, like, just not focus on, like, short-term kind of, like, growth monetization of, like, the individual developer and, like, build some of this? Because the alternative would have been, hey, all these people are using it, so we're gonna make this other, like, five bucks a month plan, monetize.
A I think, I think this might be a little bit of, like, commercial instinct, That the company has, and unclear if the commercial instinct is right. I think that right now optimizing for making money off of individual developers is probably the wrong, actually, strategy. Largely because I think individual developers can switch off of products, like, very quickly, and unless we have, like, a very large lead trying to optimize for making a lot of profit off of individual developers, it's probably something that someone else could just vaporize very quickly, and then, and then they move, they move to another product. And I'm going to say this very honestly, right? Like, when you use a product, like, Podium on the individual, on the individual side. There's not much thing, not much that prevents you to switch onto another product. I think that will change with time as the products get better and better and deeper and deeper. I constantly say this, like there's a book in business called like seven powers. And I think one of the powers that a business like ours need to have is like real switching costs. But like you first need something in the product that makes people switch on and stay on before you think about how do you make people switch off. And I think for us, We believe that there's probably much more differentiation we can derive in the enterprise by working with these large co…
AI assessment note: “optimizing for making money off of individual developers is probably the wrong, actually, strategy”
Answered raw tape
D 5 · C 4 · P 4 · Cm 4 4.30
Q Can you define Cascade because that's the first time you brought it up?
A Yeah. So Cascade is the product that is the actual agentic part of the product, right? That is capable of, of Taking information from both these human trajectories and these AI trajectories, what the human ended up doing, what the AI ended up doing, to actually propose changes and actually execute code to finally get you the final work output, right? I'll even talk about something very basic. Cascade gives you a bunch of code. We want developers to very easily be able to review this code. Okay, then we can show developers a hideous UI that they don't want to look at, and no one's going to really use this product. And we think that this is like a fundamental building block for us to make the product materially better. If people are not even willing to use the building block, where does this go?
AI assessment note: “Cascade is the product that is the actual agentic part of the product”
Answered raw tape
D 4 · C 4 · P 4 · Cm 4 4.00
Q That's true. That's true. My impression is a bunch of you are geniuses. You sit down together in a room and you get all your data. You train your model, like everything's very smooth sailing. Um, what's wrong with the image?
A Yeah. So probably a lot of it just in that a lot of our serving infrastructure was already in place before then. So like, Hey, we were able to knock off one of these boxes that I think a lot of other people maybe struggle with the open source serving. Offerings are just, I will say not great in that, in that they aren't customized to transformers and these kinds of workloads where I have high latency and I want to like batch requests and I want to batch requests while keeping latency low. But one of the weird things about generation models is they're like autoregressive, at least for the time being, they're autoregressive. So the latency for a generation is a function of the amount of tokens that you actually end up generating. Like that's like the math. And you can imagine while you're generating the tokens though, Unless you batch a lot, it's going to end up being the case that you're not going to get great flop utilization on the hardware. So there's like a bunch of trade-offs here where if you end up using something completely off the shelf, like one of these serving things, serving frameworks, you're going to end up leaving a lot of performance on the table. But for us, we were already kind of prepared to sort of do that because of our infrastructure that we had already built up. And probably the other thing to sort of note is early on, we were able to leverage open source…
AI assessment note: “there's like a bunch of trade-offs here where if you end up using something”
Answered raw tape
D 4 · C 4 · P 4 · Cm 4 4.00
Q I mean, like, your customers, you know, everyone says it's 89% on Windows, right? Like you said, you live in Windows, you will never, you would never not see something that missed
A So I think in the beginning, part of the reason why we were hesitant to do that was, like, a lot of our architectural decisions to work on across every IDE was because we built a platform agnostic way of running the system on the user's local machine that was only buildable, easily buildable on, like, on dev containers that lived on a particular type of platform, so Mac was, like, nice for that, but now there's, like, not really an excuse if it's, like, if I can also make changes to the UI and stuff like that, and yeah, WSL also exists. That's actually something that we need to add to the product. That's how early it is, uh, that we have not actually added that.
AI assessment note: “part of the reason why we were hesitant to do that was”
Answered raw tape
D 4 · C 4 · P 4 · Cm 3 3.85
Q Any thoughts on why? Like, is it, I don't know.
A I mean, emergence of intelligence. I think, I think maybe one of the things is like, we don't even know maybe like five years from now, what we're going to be running our transformers. Right. But I think it's like, we don't, we don't 100% know that that's true. I mean, there's like a lot of maybe issues with the current version of the transformers, which is like the way attention works, the attention layers work, the amount of computers quadratic in the context length. Because you're like doing like an N squared operation on the attention blocks basically. And obviously, you know, one of the things that everyone wants right now is infinite context. They want to shove as much crap as possible in here. And the current version of what a transformer looks like is maybe not ideal for that. Right. You might just end up burning a lot of flops on this when there were probably more efficient ways of doing it. So I'm, I'm sure in the future there's gonna be tweaks to this, but it is interesting that we found out interesting things of like, Hey, bigger is pretty much always better. There are probably ways of making smaller models significantly better through better data. That is like definitely true. And I think one of the cool things that the stack showed actually was they did a, like a, I think they did some ablation studies where they were like, Hey, what happens if we do, if we do dec…
AI assessment note: “I mean, emergence of intelligence. I think, I think maybe one of the things”