The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Dr. Ben Lee no published score: no usable exchanges on raw tape, and a fair score needs 8+ · coarse estimate ≈4.5/5 from 12 produced feed exchanges record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
12exchanges match
0on raw tape
0redirected or not addressed
Answered produced feed D 5 · C 5 · P 5 · Cm 5 5.00

Q like, really try to understand. The topic basically being how much of inference compute Might move from central cloud infrastructure to the edge, and then how far to the edge, of course, being another question. I think we should start by actually defining those categories a little bit. How do you think about the categorization of, like, where compute can occur? Then we'll talk about each of those categories individually.

A Right. So even before we talk about generative AI there, for classical compute, cloud computing in general, All of the services we loved and changed the way we live and work today. There are three levels generally I think about for compute. The first is massive hyperscale data centers, the ones run by Microsoft and Google and Amazon, hundreds of thousands of machines, massive facilities. That's what most people think about when they think about cloud computing. At the other end of the extreme, Would be a personal devices, consumer electronics. So you think about your phone, you think about your tablet, uh, your, your, your laptop, uh, plenty of compute can happen there as well. There is a perhaps less understood, uh, middle layer or intermediate layer called edge computing. And edge computing really means that there are times where you don't want to go all the way to this remote massive facility, uh, And wait for the data to go out to that data center and then come back. You might want to access some compute that's a little bit closer to you, maybe in the same city, maybe in the same geographic region, that's edge computing. So they're still going to supply really capable, high-performance machines, these servers, um, but you don't suffer those longer communication times or latencies that you might if you, um, if you were to go to that remote, massive data center.

AI assessment note: “There are three levels generally I think about for compute.”

Answered produced feed D 5 · C 5 · P 5 · Cm 5 5.00

Q here. But then it Seems to me that because AVs were generally delayed or maybe the need wasn't as high, like what we've got today, if you just look at the infrastructure today, it's seems like the vast, vast majority of classical compute even, um, except for stuff that's sitting in like mainframes at companies is in the cloud and the big centralized data centers. Do I have that right?

A That's right, and this is a decades long trend. I mean, we've seen this progression, uh, this adoption of cloud computing over the last 15 to 20 years, and there are a couple of reasons, uh, we are seeing that shift, or we have seen that shift. Uh, the, the first is that, uh, Computing in a massive data center run by the hyperscaler companies, these big tech companies, is much more energy efficient. They know how to deploy these facilities. They know how to cool them and build HVAC systems, um, efficiently. So they're incurring very, uh, very small overheads per watt of compute. There's this industry standard metric called, uh, power usage effectiveness or PUE. And that's the ratio of how, of the power you're using divided, compared to the power that's going to compute. So Google's PUE is close to 1.1, which is to say for every watt going to compute, there's an additional .1 watts going to the overheads of power delivery or cooling or whatever. So that's really incredibly efficient. And most mom and pop data center operators, most enterprise data center operators don't get the scale and efficiency that these hyperscalers do. Um, the scale also gives a second key advantage, which is the ability to share hardware. So you buy the hardware once, and you have lots of users sharing the same physical hardware. That allows us, again, to drive the cost down, allows the hyperscaler opera…

AI assessment note: “That's right, and this is a decades long trend.”

Answered produced feed D 5 · C 5 · P 5 · Cm 5 5.00

Q that is the next wave of applications for AI, right? And so maybe we go back to the autonomous vehicle world and things like that, where like latency making decisions in near real time does become really important. Robotics being another category that could be a major user of AI compute, but needs really, really low latency. Is that part of the argument for shifting some compute to the edge?

A Yes, absolutely. So The class of compute you mentioned autonomous vehicles, robotics fit into what we call cyber physical AI. So cyber physical systems are those that have a cyber component, a computational component, but also interact with the physical world. And once those interactions with the physical world arise, then we care about responsiveness. Because that underpins safety guarantees, and the ability to make sure that your robotic arm is able to respond quickly enough to hazards, your autonomous vehicles are able to do so. So I, I agree that there will be cases where we will need those really low latencies, and that is going to require edge computing much closer to the user, so we have much shorter internet delays, network delays.

AI assessment note: “Yes, absolutely. So The class of compute you mentioned autonomous vehicles, robotics”

Answered produced feed D 5 · C 5 · P 5 · Cm 5 5.00

Q hard to do, and indeed is, but these are all hard problems. So if that happens, Do you think that we are going to see a significant portion of that inference workload move to that type of scale? Is that the right scale? Like, should we be looking at 10 megawatt sites, a hundred megawatt sites, one megawatt sites? Like, how far to the edge do we want to go?

A Yeah, absolutely, and I, I agree with the premise of that question, 100%. I think that there are two reasons to go to smaller, many smaller data centers. The first is the one you mentioned, power, power provisioning, uh, and connections to the grid. The second is, uh, the fact that you don't need massive GPU coordination for an inference workload. Um, I, I guess the catch might be that if you are thinking about Your existing edge data centers. Maybe you've got data centers in downtown Los Angeles or something like that already serving workloads. Those workloads may not be configured to handle GPU and AI compute. Uh, they may have, uh, power delivery infrastructure that was optimized for CPUs. They might have, um, HVAC systems optimized for the much lower power density of CPUs. So it's not simply a matter of Pulling out your CPUs and replacing them with GPUs. You're going to, you may have to retrofit the, the facility itself to support that. Uh, but I, I agree. I think finding capacity there may eventually become easier than finding the next, uh, thousand megawatts.

AI assessment note: “I think that there are two reasons to go to smaller, many smaller data centers.”

Answered produced feed D 5 · C 5 · P 5 · Cm 4 4.85

Q big centralized cloud data centers, or will some or much of it potentially shift either to one of the other two categories you described, sort of edge, uh, localized or fully localized on device. So let, let's talk about the edge version first, which is essentially Smaller data centers, still data centers, but smaller and more local. What's the argument for why that might happen? And what are the limitations?

A So, so the argument in favor of edge computing is mainly the proximity to the end user, right? So when you, so we have, we have been conditioned in an era before generative AI that when we access internet-based services like a search engine, We expect the answer to come back on the order of a hundred milliseconds. That, that is the order of magnitude that we're, we're talking about. And, and as a result, to get those hundred millisecond latencies, oftentimes you require computation closer to the user. So you don't have to travel across the internet. You don't have to travel from the west coast out to the east coast and back again, uh, the data, I mean, um, and, um, and get that answer back, uh, in a timely way. What is interesting with generative AI is that we are being reconditioned to tolerate much longer delays. So if you use something like GPT or you use something like Claude or your favorite chatbot, oftentimes it's just sitting there thinking for seconds and seconds, maybe tens of seconds before it gets you the first token. So, so the question there is to what extent we care about that latency and need that really fast responsive access to the answer.

AI assessment note: “the argument in favor of edge computing is mainly the proximity to the end user”

Answered produced feed D 5 · C 5 · P 5 · Cm 4 4.85

Q not like we're waiting on workloads to show up that could accommodate this. And yet, if you look at everybody Most everybody building data centers, certainly the hyperscalers, and I think the colos and, and folks as well, you know, the focus continues to be on, we gotta find big sites for big data centers. Why don't we see more development of this small, smaller scale edge AI inference world?

A I think it really depends on, on the workload and the application, and we don't know, I, I view AI as a more fundamental basic technology, and we don't necessarily know what application or capability will be layered on, on, on top of it. I, I'd say that we've been talking about edge data centers a lot. There are other words for this, uh, type of data center. A content distribution network is one of those examples, a CDN, um, or a point of presence, uh, a BOP that, Uh, the facilities are sometimes called, and they exist in fairly significant numbers. Content distribution networks ensure that when you want to access, for example, newyorktimes.com or wsj.com, Your web page is not being served from the other end of the country. Those web pages are sitting close to you because the content distribution network took those updated web pages and moved them to facilities near you, data centers near you. Likewise, companies like Meta, um, when they have Instagram or when they have, uh, these social media applications, they also have these points of presence that supply data from, um, Local points of presence rather than retrieving content for your feed from across the country. So we already see that, but these are application level, uh, performance requirements, whether they be for social media or for other sort of news content. Once it becomes clear what applications of AI really drive f…

AI assessment note: “Once it becomes clear what applications of AI really drive further inference, uh, deployments”

Answered produced feed D 5 · C 5 · P 5 · Cm 4 4.85

Q up local, that, that's a significant shift and has, has pretty profound implications for the energy picture as well. Are you saying that 80%, just to Just to pin you down even a little bit more, is that local in the sense of being at the edge, or is that local in the sense of being on device, or like, what do you think the split ends up being there?

A Yeah, so I, I think Of the 80%, I would say most of that will be on the edge. Um, like it, I, like I suspect it is today. I, I think that, um, if you look at what we, what we talked about earlier, the content delivery networks, um, points of presence, they've probably identified 20% of the content that 80% of the people will be looking at most of the time, and they're putting it at the edge. I think maybe on the order of one percent, Ends up being put on your consumer electronics. Actually, even for today's compute, When we set aside AI, there is a trend towards, um, consumer electronics, hiding that flow of data back and forth between the, uh, the device and the edge for you. Right. So sometimes they'll like, if you use a, a cloud storage service like Dropbox, or if you're using a photo storage service, they will let you pretend that you have access to all of your videos or all of your photos and all of your documents. And they will transparently, behind the scenes, move things back and forth between the data center and your local device. So you may think you have all of it, but maybe you've only got a tiny sliver, less than one percent, on your local device.

AI assessment note: “Of the 80%, I would say most of that will be on the edge.”

Answered produced feed D 5 · C 5 · P 5 · Cm 4 4.85

Q Okay, so then on to inference. So you're saying inference does not contain that same challenge. So is there any, what is the downside to shifting inference workloads to the edge?

A Um, to my knowledge, there isn't much of a downside because the reason why inference, um, is amenable to edge computing is because when you send a prompt to, uh, for processing by a large language model, that prompt is probably handled by one GPU or maybe Eight GPUs inside a single machine. So, and, and the reason that is, is because the model sits in that machine, the data sits in that machine, and all of your prior conversations with that bot have, are sitting in that machine. And it's a very localized, uh, piece of compute that needs to be done. And you don't need tens or hundreds of GPUs to be coordinating to give you an answer back. You've got that one GPU or a tightly coupled GPUs giving you that answer back. And that is amenable. That is great for, for edge computing, and we can certainly supply that.

AI assessment note: “to my knowledge, there isn't much of a downside because”

Answered produced feed D 5 · C 5 · P 4 · Cm 4 4.60

Q you could do on-device, but, you know, you, you pull up your ChatGPT app or whatever, and of course it's going to Send you back out to the cloud, or maybe to the edge. But, you know, these other challenges of thermal management and things like that are hardware challenges. Where are we in the progression of on-device inference? Is it coming? Is it not coming? Do we not know?

A I think the assumption with on-device inference is that you'll be able to shrink the model without loss in performance for the tasks you care about. That is the primary strategy the computer scientists have been taking. On the hardware side, Uh, we have made strides in developing custom chips, custom silicon for the specific types of tensor algebra that are required for, for machine learning models. So we know how to build those chips and that gives us energy efficient compute, higher performance. Uh, we know how to build, uh, really capable memory systems or solid state disks. So when your phone now has Hundreds of gigabytes of memory on it, or hundreds of gigabytes of storage on it. So there's a question of, well, maybe you'll end up using less of it for your photos and more of it for your AI model, something like that. Um, so I think there, there are fairly significant resource constraints, but I don't think that they are insurmountable in the sense that more intelligent hardware design and more intelligent hardware management could go some ways in terms of, um, making these AI models feasible on the device.

AI assessment note: “significant resource constraints, but I don't think that they are insurmountable”

Answered produced feed D 4 · C 5 · P 5 · Cm 4 4.55

Q them or the optics or whatever it is that's communicating between them. And that for some reason that you can explain to me makes model training more effective. Um, Is there a similar dynamic in inference? Is there a technical reason why that you are, you're paying a penalty if you shift to smaller, uh, data centers at the edge? Or is there no technical reason why that's, it's suboptimal?

A Right, yeah. Let's, let's talk about the training piece first. Um, the reason why we need a thousand megawatt data centers where we have hundreds of thousands of GPEs connected so closely together is because the data sets are massive. Uh, and the models are massive. We're trying to learn, um, on the order of a trillion parameters, uh, for these machine learning models, these AI models, and we're trying to do it on the wealth of data we find in the internet. There's no way that any single GPU can handle that much data, so what we end up doing is partitioning the data into smaller pieces and then handing each GPU a slice or a partition of this data. And each GPU will churn on its own share, on its own partition of the data, and learn the models that work best for its piece of the data. And all the other GPUs in the data center are doing the same thing on their partitions of the data. Periodically, What they will do is they will compare notes. They will share the weights that they've learned. And this sharing is really, really expensive. And some of the people in the energy space may know that there are massive energy fluctuations or power fluctuations we will see in data center usage when the GPUs go from this computational intensive phase where you're learning the model weights to this communication intensive phase where they're comparing notes and sharing their intermediate res…

AI assessment note: “For inference, we don't see that effect.”

Answered produced feed D 4 · C 5 · P 4 · Cm 4 4.30

Q Instead of hundreds or, or thousands of megawatt data centers. Does it look similar? Is that you have a bunch of, a small number of regions that, um, have like a really high concentration of those 15 to 15 megawatt data centers? Or could it be much more dispersed because the whole point of this is really low latency and local and you don't need them to be as clustered?

A I think there are lots of different aspects at play in, in terms of Uh, data center siting. I, I think the redundancy is definitely one of them. And I have trouble disentangling the role that some of these other factors play as well. Some people talk about tax breaks and incentives from local companies and local states. Uh, some people talk about proximity to internet exchange points. So not only are you talking about, uh, congestion-free power movement, but you're also talking about congestion-free data movement into and out of the data centers. Northern Virginia is, Has that. Um, and then of course the availability of the power itself. Um, I, I guess I would say that when you start talking about many of these smaller data centers, from a redundancy perspective, it might be okay that they're not all geographically clustered as long as you have a strategy for rolling over the compute or rolling over the workload to spare capacity somewhere within that region. That has a similar performance profile or some sort of similar, uh, latency or delay characteristic. Um, and so that's really, that's really the concern, whether you have robust and geographical redundancy and, and resilience there.

AI assessment note: “it might be okay that they're not all geographically clustered as long as”

Answered produced feed D 4 · C 4 · P 4 · Cm 4 4.00

Q trying to think of why you wouldn't do that. Um, you know, you need to sort of house all of the, you need to have a fair amount of memory, you need to house all the model weights and so on in every individual data center if you're going to do that at the edge, right? So is there, there's got, got to be some minimum viable scale, I assume.

A Right. And maybe to give you a sense of the type of data centers we were talking about in the past, um, Again, in a study that we had done with Meta, we looked at 15 of their data centers before generative AI, and the scale of those facilities were somewhere between 15 to 50 megawatts, right? So less than a hundred megawatts, and certainly that is, that was fairly, fairly conventional, uncontroversial to build those sites of data centers in the past. Um, so that, that's the starting point, I think, in terms of The, the scale. Now, as you scale down towards, for example, one megawatt, uh, not clear, uh, at what point things, uh, start making less sense.

AI assessment note: “as you scale down towards, for example, one megawatt, uh, not clear”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.