The Exchanges, every show

Every argument clarity score on this site is built from rows on this page, here across all 44 shows. Each question and answer was assessed with names hidden, the hosts' own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

shows every show 44 of 44
every show
Stephen Balaban no published score: a fair score needs 8 or more exchanges on raw tape on one show record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score rests on one show's raw tape, the show with the most assessed exchanges, and shrinks small samples toward that show's cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
14exchanges match on 44 shows
14on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q Yeah. That's the highly optimized, uh, storage. What else? The networking part and what other pieces?

A So, so I, I, I was talking about this one cloud cluster product that We've got, and the, the way of just for everybody to think about this is like, okay, well, look, you've got a bunch of GPs. Let's say you've got a cluster of 10,000 GPs. Well, I want to partition that cluster up, and so what it is, is it's a bunch of GPUs, some CPU servers as well, because you need to have an orchestration, uh, fleet as well, and then you've got some storage, and, um, all of the CPU servers and the storage servers and the GPU servers are interconnected with the, the storage, so they can quickly read and write from it, and, um, so there's And that, that communication happens over what's called, you know, the in-band network. And then there's the compute fabric, which is where I was talking about where all of the sort of weights and, uh, feature activations are being shared, uh, throughout that compute fabric. And then there's an out-of-band monitoring network where you've got access to whether it's BMC or, uh, some of your DPUs. Um, and When you are trying to create a sub-partition of a 10,000 GPU cluster, you need to simultaneously partition the in-band, the out-of-band, and the compute fabric. Okay, so, like, that complex coordination between we've got a bunch of bare metal systems to, hey, we've got a virtualized system that has, you know, what's called RDMA, you know, RDMA, remote direct me…

AI assessment note: “simultaneously partition the in-band, the out-of-band, and the compute fabric.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q How real is the movement against data centers from the global community and how do you think about, uh, how to respond to it?

A Well, it's certainly, it's like very popular in the news right now. I'd say that, um, it's definitely very real. I mean, I think that rightfully communities that host any type of large capital project, whether it's a power plant or a, uh, solar farm or a data center or a distribution center, right? Those communities want to have a seat at the table. I'd say in general though, I spend a lot of time reading through a lot of the comments from communities, and people want jobs. They want tax revenue. Any major capital development Is going to bring a lot of tax revenue and it's going to bring a lot of jobs and it's going to bring investment into their community. And what they really are voicing, I think is one is having a seat at the table while, while this stuff is, you know, being developed. I think that that's an important thing is just to have their voices heard and that, that the developers coming in and actually understanding the community. The other thing to kind of, I think, keep in mind is That there's a lot of misinformation out there. So for example, Every single modern deployment of, let's say, a Blackwell class or a Rubin class GPU, you know, the VR, GBN VR GPUs. Um, these are oftentimes in a closed direct-to-chip liquid cooling system that's connected to a dry cooler, which means that there's almost zero evaporation. It's not using evaporative cooling. It's using a dry…

AI assessment note: “I'd say that, um, it's definitely very real.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Super helpful. If two companies have the same chip fundamentally, how do they extract more value from it? What, what needs to happen to maximize the Usefulness of that chip.

A If you look at the cost structure of let's say one GPU hour of time, you know, we were talking about H 100. The largest part of that cost structure is the depreciation that is associated with that GPU hour. And, um, basically you can think of a utilization metric as being like kind of a multiplicative factor on that. So one over the utilization. So if you, If you use your capital asset, 50% of the time, you will have on a per hour basis, twice one over 0.5, the amount of per hour depreciation expense associated with that. And so I think that the number one way that companies are, you know, sort of gaining a unique advantage is, well, how can I build a cloud product that is beloved by people that is going to drive a high utilization? And, um, you know, in addition to that, the market, as we mentioned earlier, for on-demand compute basically The retail pricing is obviously much higher than the wholesale pricing. So the retail is like on demand, spin up a GPU, spin down a GPU, normal cloud service. The wholesale is sort of buying 10,000 GPUs for five years, for example. And so one of the things that we do at Lambda is really try to figure out, hey, how can we sort of get the most dollar utilization and percentage utilization out of The capital deployments that we do. And that's, that's by making great cloud software that makes it easy for somebody to spin it up and down. So for ex…

AI assessment note: “number one way that companies are, you know, sort of gaining a unique advantage”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Do you need for performance reasons to be close to the customer the way you, you need to have regions in cloud?

A You know, it's super interesting. A lot. I get this question a lot and people they're like, well, does latency matter? Does, so I'll tell you what, what matters and what doesn't matter. You can look at your own utilization of whether it's ChatGPT or Claude or Grok or Gemini, and you can see, hey, a lot of the things that I'm doing, I kind of shoot it off. I come back later and there's a research report for me. Maybe it's a long running agent workflow. In those cases, latency doesn't matter at all. The only thing that matters is your cost per token. That's all that matters. And, um, So that's been a really interesting change. I think that, you know, the old school traditional legacy cloud business was so latency focused because of some of the applications, but this new fleet of AI applications are far less latency sensitive. So that's one, but there is the caveat, which is this governance and data governance is becoming an important thing. And a lot of countries are wanting to have the AI compute that their citizens are using Be run out of their own country so that they can, you know, at least have their own, their perception of control or whatever. And the, you know, that is, that is another, that is an element to it, but I'd say that from the latency, there's no technical reasons.

AI assessment note: “from the latency, there's no technical reasons.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Are some of the people that were there at the beginning still around? I think you started the company with your brother. Is that right? And your brother is still at the company.

A Yeah. And so in terms of like the early people, um, basically It's not, I mean, not even basically of the four people who are making DreamScope, me, Michael Balaban, my co-founder and fraternal twin brother, Chuan Li, who's our chief scientific officer, and then Steve Clarkson, who's, um, an engineering leader at the company and, you know, has a bunch of folks reporting into him. Uh, now, you know, they're all still at the company. Um, the next hire, one of those, uh, gentlemen named Mitesh, Uh, Agrawal, who's one of the, the next hires in that team. Um, he was with the company for maybe eight years or something like this. Um, yeah, something like eight years. And, uh, then he eventually, uh, left and, and joined another former Lambda team member, Thomas Summers to start Positron, which is, um, uh, an accelerator company. And they're like now valued at over a billion dollars. And, uh, so, uh, not only has like the original team stuck around, but we've already started to kind of see what like a Lambda, uh, alumni, a Lambda mafia network looks like in, in the world, Lambda lab member alumni.

AI assessment note: “they're all still at the company.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Do you worry about model training and model inference becoming, I don't know, 10 X more compute efficient and what that would mean in terms of the buildup?

A I think that generally speaking, What you're seeing is if, let's say you do become 10 times more efficient, I think that that just means that everybody is able to process 10 times more tokens, and there's, there's still the same fixed amount of compute in the world at any given point in time, and so in the early days, it's funny, we used to talk a lot about this back in, let's say, 2017. Oh, well, maybe there's gonna be some new type of model, let's say, that will look more like a random forest model, which The audience might some, some members of the audience might know you can kind of train a random forest model on a MacBook, right? And there was, there was this concern that was kind of persistently raised around like, well, okay, what happens if you have this sort of like, um, adjacent disruption on the model side of things? And so far we haven't seen that. And again, everything that we're building towards is, Sort of based on these scaling laws, which is really about scaling up this architecture. So, um, I don't really foresee a very likely outcome where we have this huge model disruption that would cause a decline in the demand for compute.

AI assessment note: “I don't really foresee a very likely outcome where we have this huge model disruption”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q But there is some element of, uh, computerization, right? The price of rental of a GPU is going down. But, uh, what you're saying is that to some extent, it doesn't matter because it's only one layer of the cake.

A Yeah, so when you look at, for example, I think it's actually worth doing is to try to, like, kind of dig into some of the methodology on, for example, an index like, there's, there's the, there's the index that's on Bloomberg for H 100 rental prices. What we're actually seeing in the market is that, first of all, there's two different rates. There's a public cloud on demand rate, and then there's a long-term rental rate. And I think that some of these indices don't properly take that into account because what we're actually seeing is a very consistent, if not increasing long-term rental rate. And Very consistent and increasing on demand rental rates. And so what happens is if, if, if the index mix, for example, if the methodology and the index biases towards long-term contracts being a bigger part of the volume, that will look like a decline in the index when the reality is it's just a decline in the mix of the index. This is method, you know, the index is covering.

AI assessment note: “what we're actually seeing is a very consistent, if not increasing long-term rental rate”

Answered raw tape D 4 · C 5 · P 4 · Cm 4 4.30

Q Do you think that today or in the near future, we're going to be in a multi-silicon kind of world? Is there like room for different players beyond Nvidia?

A Well, I mean, I think that we're already in a world where there's a huge amount of competition from massive, massive multi-trillion dollar companies, and they're all trying to fight for the same thing, which is to be the best chip in the world for running and training Neural networks, essentially. NVIDIA's built a great product that has gotten a lot of distribution and has a great platform of developers who love what they do, and you have to take into account not just the cost of the chip, right? The price of the chip is one aspect, but, you know, you have to take into account the entire software ecosystem and what's been developed. So one of the big people talk about what's NVIDIA's moat. One of the big moats they've got is just The QDNN stack. It's not just CUDA. It's, you know, CUDA is sure. That's like the water we all swim, but like CUDNN has got so many, you know, matrix multiplication, routine optimizations baked into it.

AI assessment note: “we're already in a world where there's a huge amount of competition”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q Let's talk about the financing stack. So presumably it's a commission of equity and debt. How does it all work?

A Yeah. So, um, the way that it works is that you You know, you could really fragment it into this two parts, which is like financing your on-demand cloud versus financing an off-take agreement, which is like a longer term commitment. And on the on-demand cloud, you're kind of looking at Lambda's credit quality. On the off-take agreement, you're kind of looking at the credit quality of the end customer who's paying the bill. And so what you do is you just, you know, take, uh, your offtake agreement, you take this chunk of GPUs that you're deploying, you take a lease or the, the property and you kind of put it into a box and you can go to the private credit markets and you can come up with, you know, an asset based loan. You can, you can get a, a variety. There's a variety of different methodologies, uh, for financing it. Um, most of it is just some sort of like, Special purpose vehicle that's designed to finance this particular deployment. With a very known and easy to underwrite, which is basically just a fancy way of saying the, you know, finance, uh, term for just understood assessing the risks and the downsides of a particular credit investment. And, uh, there's, there's a, there's a vibrant private and, uh, you know, there's a, there's a, there's a vibrant credit market for that on the on demand cloud side of things. It's, you know, Not quite as mature as when there's a, for…

AI assessment note: “most of it is just some sort of like Special purpose vehicle”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q You had a quote where you said that AI won't write software, it will become the software. What do you mean by that?

A So, uh, that's in my sort of like idea around what I call neural software. And, or, you know, a neural computer, neural operating systems. And the best way to kind of get this experience is to go to your ChatGPT or your Claude and say, hey, just, um, you know, render for me an ASCII art desktop interface, ok? So you're working in purely in the domain of text, and I want you to just pretend to be an operating system for me. I'm gonna say, click on this, You know, open up this, and I want you to just behave like a computer. So give it that prompt, ok? And, um, what you're gonna see, I think, is that you're gonna see that, uh, that sort of future of the large language model becoming the software and not generating the software. And, um, this results in an extremely sort of squishy and flexible Way of interfacing with a computer where it's not possible to have a bug, only a misunderstanding about the prompt and what you've asked for. And I think that for a lot of the pieces of software on your computer, you might see that taking over where, you know, you can get the glimpse of the future with this ASCII art, and then eventually it'll also have a multimodal network that's generating every pixel on your screen. As well as every audio waveform that comes out of your speakers. The advantage to this is that you can really sort of dream up software. The only the part that is being experi…

AI assessment note: “large language model becoming the software and not generating the software”

Answered raw tape D 4 · C 4 · P 5 · Cm 4 4.25

Q Maybe quickly just Go back to the very origin, because I think you've been in the effectively in the AI world the whole time, but are coming from a very different angle, uh, with multiple pivots. What did you start with and when?

A Well, you know, with the complexity of the business, you can now see, you know, the complexity, the capital intensivity, just the sort of not fitting into a box. And you can see why we've oftentimes not had a lot of traditional venture investors in, in Lambda. And, uh, you know, our, all of our investors have done exceptionally well, but, but they, they've kind of come from more often than not outside of traditional, let's say mainline Silicon Valley VCs. And, um, so Just going back to the origin story. I started Lambda in 2012, and we were a facial recognition software company. So I was training convolutional neural networks to do face and image recognition, and we eventually hosted that on API. I was training those confidence on a four X NVIDIA, uh, uh, GTX five 80 workstation that I, uh, that I bought from a friend who had built it actually. And, um, And this was, you know, really pretty avant-garde stuff at the time. Most people didn't really believe in what was called the field called deep learning at the time.

AI assessment note: “I started Lambda in 2012, and we were a facial recognition software company.”

Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q And what you described for Frontier in France is that conceptually the same thing as what happens for training, this concept of just distributing a task massively across a bunch of GPUs. What happens during a training run from a compute standpoint?

A Generally speaking, when you're doing a training run, you might think that you might be some sort of split between the backwards pass and the forward pass on the model. And the backwards pass might be, let's say, two thirds or more of the compute and the forward pass, which is basically the same thing as inferencing, uh, is, you know, the remainder. And one of the realizations that I think has been made over the last bit of time is that the type of infrastructure that you'd want for, uh, Doing a large-scale training run can be reused to do the inferencing of that model. And, um, What I mean by the sort of frontier inference and the fact that the inferencing is being done in a distributed way, you know, you'll have like a mixture of experts model and there'll be a different, basically starting strategies for how you put those experts onto different servers and to different GPUs. Um, and you know, the models can be very large. They may not fit on one single, uh, Rack. Or, you know, they may not fit on one single server. They, they might, they might need to be distributed across Different servers to even just do the forward inference pass. And so that's where sort of distributed frontier inference kind of comes into the picture, right? Because like, if you're doing a small model, let's say Llama, that the users might be, you know, familiar with, or, uh, some of the quantized small…

AI assessment note: “the backwards pass might be, let's say, two thirds or more of the compute”

Answered raw tape D 4 · C 4 · P 4 · Cm 3 3.85

Q You also talked about one person, one GPU. Is that your, your vision for the future? Unpack that for us.

A So, um, you know, before people really believed in the AI thesis, I, you know, when I was pitching our series B and C, I would kind of talk a lot about the similarities between, let's say the computer industry and the AI industry. I really felt like AI was forming a set of generational companies. Um, and there was going to be a set of generational companies that got minted with the changes that were coming with AI. And this is like in 20, 20, 21. And if you read about the history of Apple, for example, in the early days, the motto and the, the sort of the credo at Apple was one person, one computer, one person, one computer. And, you know, there's a sense of humility that's embedded in this one person, one GPU, which is the one person, one computer. You think about how visionary Steve Jobs was. That was, you know, Apple was, you know, what founded in 1976 or something like this. The Macintosh came out in 1985, or 1984, excuse me. 1984, you know, whatever. Um, eight years or so after founding. Is that one person, one computer yet? No, not even close. Alright, so 1984 to 1994. Alright, well, is it one person, one computer? Well, we're just starting to have the internet boom, so we're, we're, we're, we're, I mean, not quite there yet. 2004. We finally have broadband internet access, and maybe for the first time in the United States, there's not quite one person, one computer, but …

AI assessment note: “there's a sense of humility that's embedded in this one person, one GPU”

Partly raw tape D 3 · C 4 · P 4 · Cm 3 3.55

Q And when, uh, we think about compute costs, what, what costs the most money? Is that model size at memory bandwidth? Is that latency? Uh, does context window and like those very, very large context window, do they change anything to the compute cost? What, what costs the most money?

A As I mentioned, like the biggest component of the unit costs for a cloud service like this is the depreciation expense. And, um, within that, you know, is basically some sort of bill of materials for the servers that are in the data center, which is by far and away the biggest portion of the cost. If you were to talk about the capital stack, let's say you can go back down to power generation, two to three million dollars a megawatt, two to three billion dollars a gigawatt for a power plant. Um, the data center is between 10 and fifteen billion dollars a gigawatt for, um, building the data center. And then the compute, the servers can be anywhere from 35 to forty-five billion dollars a gigawatt. And, um, Within that, that's so, so you can see the server portion is obviously by far and away the largest, and that's like a big part of the depreciation expense. And then the, um, within that, obviously you have, um, The sort of server and cluster bill of materials, which is primarily the GPUs. Um, if you were to kind of break down NVIDIA's, uh, bill of materials, then, you know, you can kind of get better allocation towards, uh, where those costs are coming from. But certainly in the most recent period of time, memory expenses, you know, memory has gone up a lot in price and, uh, you know, there's, there's very few vendors, right. You know, for HBM memory with Samsung, Hynix,

AI assessment note: “biggest component of the unit costs for a cloud service like this is the depreciation expense”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.