The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Jared Quincy Davis no published score: only 6 usable exchanges on raw tape, and a fair score needs 8+ · coarse estimate ≈4.0/5 from 6 raw tape exchanges record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
6exchanges match
6on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q And do you think that's just like a quality and quality control issue for those systems? Or do you think it's just some form of complexity with some failure rate per component that's rising? Or like, what do you think is sort of the driver of that?

A I think it's more of the complexity has grown, right? And we're in a different regime now, you know, I think that it's fair to say, so maybe stepping back again to definitions, we throw the term large around a lot in the ecosystem. I guess one question is what does large mean? And one useful definition of large that I think roughly corresponds to what people mean when they invoke the term is that a large language model, you enter the large regime when the, essentially the amount of compute necessary to contain even just the model weights starts to exceed The capacity of even a state-of-the-art single GPU or single node. I think it's fair to say you're in the large regime when you need multiple, you know, state-of-the-art servers from NVIDIA or from someone else to even just contain the model, you know, just run the, basically run the training or definitely even just to contain the model. That's definitely the large regime. And so the key characteristic of the large regime is that you have to somehow orchestrate a cluster of GPUs To perform a single synchronized calculation, right? And so it becomes a bit of a distributed systems problem. I think that's one way of characterizing the large regime. Now, a consequence of that is that you have many components that are all kind of collaborating to perform a single calculation, and so any one of these components failing can actually p…

AI assessment note: “I think it's more of the complexity has grown, right?”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q And Jared, what, uh, for, for anybody who, uh, hasn't heard of Foundry yet, what is the product offering?

A Yeah. So Foundry, we're essentially a public cloud built specifically for AI. And what we've tried to do is really re-imagine All of the systems undergirding what we call the cloud, end to end, from first principles for AI workloads. And I'm sure I do this in a bit of a new way. I think the AI offerings From the existing major public clouds and kind of some new GPU clouds haven't really re-envisioned things. And by thinking about a lot of these things a bit anew, we've been able to improve the economics by, in my case, it's 12 to 20 X, um, over lower tech GPU clouds and existing public clouds. And, you know, we'll, partially based on some of these products that we'll talk about today, um, that we're releasing and a lot of new things that we're working on, um, we think we can push that quite a bit further as well. Um, and so, you know, our, our primary products are essentially infrastructure as a service, so our customers come to us for elastic and really economically viable access to state-of-the-art systems, um, and also a lot of tools to make leveraging those systems really seamless and easy, and we've invested quite a bit in things like reliability, security, elasticity, and just the core price performance.

AI assessment note: “we're essentially a public cloud built specifically for AI”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q one other thing that'd be great to cover, I know we only have a couple minutes left of your time, um, is, uh, the recent paper that you authored, um, which I thought was super interesting around compound AI system design and sort of, uh, related topics to that. So would you mind telling us a little bit about that paper and sort of what Uh, what you all showed?

A So I think it's kind of in this regime that we were just talking about where more and more often to kind of go beyond the capabilities on Frontier accessible to today's state-of-the-art models and kind of get GPT-V or GPT-VI early. Practitioners are starting to do these things oftentimes implicitly where they'll call the current state-of-the-art model many, many times. There are many scenarios where maybe you're willing to expend a bit of a higher budget. Maybe it's code or something, and if I said that I can give you a 10% better model, you know, for code, many developers might pay 10 X for that, access to that. Instead of 20 dollars a month, they might be very willing to pay 200 dollars a month, right, for obvious reasons. Um, and so there's a question of what do you do in that setting. And so people are, you know, if you're willing to call the model many times, you can compose those many calls into almost a network of network calls. Right? And, you know, I guess one of the questions is, how then should you compose these networks of networks? What principles should guide their architecture? We kind of know how to construct neural networks, but we haven't yet elucidated the principles for how to construct networks of networks, so to speak. Um, these compound data systems, so to speak, where you have many, many calls, maybe external components. And so one principle that we Star…

AI assessment note: “one principle that we Start to explore was... you can look at how verifiable the problem is”

Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q else. It could be a hobbyist. It could be a research lab. It could be somebody with just, you know, some GPUs that they're messing around with. I'm sort of curious for each one of those types of, or categories of users, like what, what is the likely utilization rate and how much more do you think it could be optimized? 10%? Is it 50%? Like, I'm just very curious.

A One of the most, I'd say, positive cases with the highest utilization, which is the case where you're running kind of an end-to-end pre-training job, right? And so that's a case where you've done a lot of work up front, you've designated a time that you're going to run this pre-training workload for, and you're really trying to get the most utilization out of it. And for a lot of companies, utilization, even during this phase, you know, is sub-eighty percent. So why? One reason is actually, That these GPUs, particularly the newer ones, actually do fail a lot, um, as, you know, black practitioners would know. And so one of the consequences of that is that it's very common now to hold aside 10 to 20% minimum of the GPUs that a team has as buffer, as healing buffer, in case of a failure so you can slot something else in to keep the training workload running, right? And so even for a lot of the more sophisticated orgs running large pre-trainings at scale, The utilization sub-eighty percent, sometimes less than 50%, actually, depending on how bad of a batch, uh, they have and the frequency of failure in the cluster. Um, and so even in that case, now, there are often, though, also large gaps and intermissions between training workloads, um, even if you have the, the GPUs are dedicated to a specific entity, you know, and so even in those most conservative cases, which all come back to…

AI assessment note: “for a lot of companies, utilization, even during this phase, you know, is sub-eighty percent.”

Answered raw tape D 3 · C 4 · P 4 · Cm 3 3.55

Q Uh, you guys have some new releases, um, as of I think a day or two ago, um, uh, can you describe what's just come out from Foundry?

A So I think right now, AI cloud is, the AI cloud business is very much like a parking lot business, and that sounds really funny because cloud is supposed to be high tech, um, and you can hardly conceive of a less sophisticated business, at least on the surface, than parking lots, and what do I mean by that? Well, there's fundamentally, there are two models in the parking lot business. One is pay-as-you-go, For pay as you go, the race is curious. You know, you may or may not find a space that you, I'm sure many of us have the experience of driving through SF and seeing a lot, a lot full sign, you know, for lot after lot as we drive around trying to park. Um, and if you do get a spot, you might pay, you know, the 12 dollar an hour, you know, rate or something like that. I'm choosing that rate because it's the rate of AWS, um, you know, for on demand. On the other hand, if you want to kind of guarantee that you'll have a spot and also have a better rate, you can basically buy a spot reserved. Um, so you can have your own reserve parking spot in your building. Maybe you pay, you know, four dollars an hour, um, so you're getting a massive discount, but it's three K a month. Effectively. Right? Which is actually pretty substantial, and if you're only using it 40 hours a week when you're in the office as a typical worker, it actually might be effectively 16 dollars an hour, um, as opp…

AI assessment note: “one of the things that we want to do with a couple of these products”

Partly raw tape D 2 · C 4 · P 4 · Cm 3 3.25

Q Yeah, do you want to break that down for sort of our listeners in terms of what you view as the differences?

A First, I guess, as a little bit of context for people, the cloud as we currently know it Is arguably one of the, you know, most important business categories in the world. That's, I think, pretty clear. The biggest companies in the world are either clouds, Azure, AWS, GCP, um, you know, core component of the biggest companies in the world, or NVIDIA, who sells for the clouds, obviously. Um, so it's clearly an important category. AWS is arguably a trillion dollars plus of market cap if you broke it out from Amazon. Um, so, you know, at one point, relatively recently, it was all of Amazon's profit and more. So it's an important category, to say the least. Cloud, as we know today, really started in Which is when Amazon, you know, dedicated to 50 people initially, I believe, to start working on AWS. Um, and they worked on this for three years with an Amazon, launched in 2006 in March with S three. Later that year, in September, they launched EC two, um, and that kind of was the beginning of the cloud. It took quite a while for this model to catch on, and people were not quite clear on why this would be useful, um, even until quite recently. And so in 2009, 2010, you know, Matei, um, Sarah, who you mentioned, um, Wrote this paper called Above the Clouds with some collaborators at Berkeley, at Berkeley Beyond Cloud Computing. And they talked about why cloud would be a big deal, and I…

AI assessment note: “First, I guess, as a little bit of context for people”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.