Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered produced feed
D 5 · C 5 · P 5 · Cm 4 4.85
Q two, is there any argument for that changing in the future? Because that reliability requirement causes so much challenge and capex, right? Why is it such a problem that we have lead times on, on gas generators, all this kind of stuff? It is because of the reliability requirement. So is it intrinsic to something about what you're doing, or is it just a function of how the Businesses evolved.
A Yeah, a fantastic question. And I think that if I were to probably, um, send one message here is no, it is not intrinsic and we should be thinking about lower reliability, uh, power delivery, uh, overall. I'll tell you, uh, why it has been, but I think that, um, I'll also get to why it has changed substantially. So for most modern software services, the compute is actually a relatively small fraction of your costs. And so now it makes sense to, um, over-provision it. You want to have 99 point actually nine nine nine percent, five nines reliability for your software services. You don't need quite that, but many of our data centers aim for four nines of, I mean, minutes of downtime a year maximum, which as you said, uh, has a large amount of costs associated with it. Now, if you think about it though, as of now, given how constrained resources are, And how costly they are, a much larger fraction of your overall service cost is in the compute. So if you went to your internal customers, if I were to go to my internal customers and said, would you rather have four nines of availability and half the capacity, or two nines of availability and twice the capacity, which do you pick? Very often, not always, very often they'll say, oh my gosh, give me two extra capacity. And if I need to have a nine, nine percent, a nine, nine percent sounds good. You all know the math. That's 3.65 days o…
AI assessment note: “no, it is not intrinsic and we should be thinking about lower reliability”
Answered produced feed
D 5 · C 5 · P 4 · Cm 4 4.60
Q few gigawatts for training, you move on to the next few gigawatts for training of whatever the next TPU or GPU is. Um, but is that enough to serve the booming? Like, I think the assumption has been, look, we're, we're training now, but that is going to result in the inference demand shooting upward. Yeah. Right. And so then that would imply it's not nearly going to be enough.
A Exactly. And so this is, and I think we're at that transition point. I mean, we, uh, said last year that, uh, we're entering the age of inference. I think with, uh, agents exploding today, that's well, well, uh, happening. So probably, I mean, the analogy I would use is, uh, from Google's early days, uh, with web search. It used to be that most of the compute at Google was dedicated to building the search index. Pretty quickly, you hoped, and unfortunately not to be true, that most of the capacity needed to be used to serve that index. Same thing here. Most of our capacity, maybe earlier on, was used for building the model, but you would hope that it transitions to serving the model pretty quickly, and you're absolutely right that we're there. So I do think that, uh, over time, also, as the efficiency and latency of these models, uh, improves, more disparate deployments is gonna, are gonna be valuable. So what I mean by that is today, Each individual token that is generated by the model takes a reasonable amount of latency, so much so that actually you might not be able to tell the difference here, let's say in San Francisco, if you're accessing content on the East Coast, maybe even Europe sometimes, relative to San Francisco. In general, for let's say maps or search or ads, that's not true. The computing is sufficiently efficient and the latency is sufficiently low that you wi…
AI assessment note: “Exactly. And so this is, and I think we're at that transition point.”
Answered produced feed
D 5 · C 5 · P 4 · Cm 4 4.60
Q for, for smaller, uh, pixel sizes making sense for data centers, does it end up being easier, um, In three years, five years, something like that, to go build a new gigawatt data center and find a site on the grid that you can interconnect the gigawatt data center, or does it become easier and or faster to build, you know, 50, 20 megawatt data centers or something like that?
A Yeah, that's a good question. In general, we found over the years that it's easier to build a smaller number of larger sites. There's still asterisks there. You don't want to be too concentrated, again, from a fault tolerance and geographic locality perspective. In other words, the argument of build as big a site as you can in one place breaks down rather quickly, but having a thousand each with .1% of your capacity has other overheads associated with it in terms of management. So I think that's Uh, it'll really come down to geographic locality and probably a medium number of medium-sized data centers. Sorry for the, um, uh, whatever, um, uh, lack of precision there, but medium number of medium-sized data centers augmented with a small number of large data centers.
AI assessment note: “it's easier to build a smaller number of larger sites.”
Answered produced feed
D 5 · C 5 · P 4 · Cm 4 4.60
Q I was going to say, why? Is it reliability?
A It is in the end provisioning for a given level of reliability. If you're behind the meter, you're going to have to do all that provisioning yourself. Now, an aspect of this that's actually quite powerful for us, and to give an example, going back to the reliability question, Uh, in March, we actually hit a significant milestone in agreements with utilities for a gigawatt of demand response across our fleet. What does demand response mean? It means that for the utility, for the one week of the year where they have maximum demand, we're willing to brown down. And that also goes to the availability commitment that we make to our customers. Why? Because that allows them to provision not for their worst, coldest, hottest, whatever it is, uh, week of the year, But then provision for the 90, whatever it is, eighth percentile, and we'll give up that capacity in exchange for, well, uh, in the end, more availability of power, less cost, both for us, but also for the ratepayers, uh, in the region. So now, if we have to do that, all the reliability work ourselves, rather than being able to shift capacity back and forth when we're not using it, like, let's say that we actually have behind the meter power generation, and we, uh, will behind the meter in quotes, What if we can, when we're not using it, give it back to the utility? In general, the way we look at it is, we like behind the mete…
AI assessment note: “It is in the end provisioning for a given level of reliability.”
Answered produced feed
D 4 · C 5 · P 5 · Cm 4 4.55
Q purposes, scale is really important. This is why we're getting these huge data centers. But for inference, I've heard Mixed things. As we shift more into inference world, it may or may not be true that you need that level of individual scale. So, in your mind, how much does scale matter when it comes to inference computing? When I say scale, I mean scale of the individual data center.
A Yeah, it's a great question, and I think you have it spot on. I remember when Google announced its first data center in Oregon, the Dallas. This was 2304 years ago, before I was at Google. Uh, 10 megawatts. And people were just stunned that a little company would go build a 10 megawatt data center. That was a big number. And actually no one else was building data centers for their own computer infrastructure at the time. And it's just grown from there. A hundred megawatts, gigawatt, et cetera. It's a really good question in terms of the split between training and serving. And so here's where to me gets perhaps most interesting. At a scale that we're operating, We want the latest, greatest, most efficient, most capable training cluster essentially on an annual basis. If you look at our announcements for TPUs, NVIDIA's announcements for GPUs, the latest, greatest is coming out every year. And every year, the latest, greatest is, by definition, better than last year. Let's pick this gigawatt number. Let's say you buy the latest greatest and you put a gigawatt somewhere, and maybe you put a couple of these down. After a few years, one, two, probably not much more than that, whoever is doing the training is going to want the new latest greatest, and then they're going to want a gigawatt somewhere else. Now you've got a gigawatt of capacity that used to be used for training. What are…
AI assessment note: “could you get away with lower scale? Yes, absolutely.”
Answered produced feed
D 4 · C 4 · P 4 · Cm 3 3.85
Q gonna be important. Where do you see the biggest opportunities? Like, if you think out into the future, how do you turn, if you were to build the same amount of capacity in five years in megawatts as you are today, Is there a world where you turn that 175 hundred eighty five billion dollars into a hundred billion dollars, and what are the things that could get you there?
A We're looking at this all the time. I mean, in other words, it is probably one of the biggest focus areas in my team. Um, I won't say biggest, but it's top three for sure. It might, might be biggest. So in other words, when we say we're spending X dollars, um, we're, we're saying that if we had had to have done this work last year, we would have to have spent 1.2 X making, making up the number. Don't, don't take it as, uh, in other words, every year we're looking to deliver substantial efficiency such that if we had to do it again, it would be way more efficient. This starts with software, and again, a lot of opportunity on the software side, but lots of opportunity on the hardware side. Let me give a very simple example. What is the ratio of power to space in your data center? In other words, if you have, let me pick a number, a hundred megawatts. How big a building do you build? And how big a building do you build for a 25 year lifetime of that building? And not just one generation of TPU or GPU or whatever, but maybe five or six generations of them. Now you could be conservative and build an infinitely sized whatever it is building and say, okay, whatever comes next, I'm going to be set. Or maybe I have to, now if you think about the watts per linear foot of a disc rack versus a GPU rack, Radically different. Like, I don't know, a hundred x different between disk and GPU. Wh…
AI assessment note: “starts with software... lots of opportunity on the hardware side. Let me give a very simple example.”
Answered produced feed
D 4 · C 4 · P 3 · Cm 3 3.60
Q All right, I'm gonna, I'm gonna force you to answer the question in a different way. You're supposed to spend whatever it is, a 175, a hundred eighty-five billion dollars this year building out new infrastructure. Uh, if you woke up tomorrow and Sundar said, you know, you gotta spend 300 now, uh, what would you go try to solve?
A You know, I think, I, I know that I, I'm not trying to dodge the question, but I very sincerely feel that actually we'd have to go scale all of them, and that every single one of those is at the limit of what we can do for the envelope that we have. Is one of them inherently easier to scale than the others among the, uh, options? Honestly, no. Uh, all, all three of those are major, major issues for us. I'm sure that there is an answer. But I'm not relaxed about any of them. This is the, uh, is a real thing. I couldn't, I couldn't pick one. I would say Sundar, wow, uh, 300. Okay, I'll get back to you as to what the exact issues are going to be.
AI assessment note: “I very sincerely feel that actually we'd have to go scale all of them”
Not addressed produced feed
D 1 · C 5 · P 4 · Cm 3 3.25
Q part of what's required is that you have to differentiate amongst the workloads such that some can operate as necessary at really low latency and others at higher latency. You know, Google can kind of do all that in-house. I mean, you have customers for Gemini, so you have to serve those customers, but you have more capability than, than most. How do you think it disseminates out beyond Google?
A Um, so I think that, uh, it's, it's a good question. It's something that we think a lot, a lot about. In other words, what we want to do is we want to design end-to-end systems that, taken together, create capabilities. This word capability is actually essential to what we discuss internally a lot, so I appreciate the question. Create capabilities that otherwise wouldn't be possible. And I do think that it comes down to this vertical integration. In other words, for us, for let's say our TPUs, we co-design them with a building. We co-design them with the power generation source. We co-design them with the DeepMind team that builds Gemini models. So it's the software above, the models above that, the chip design, which we do in my team as well, that's integrated with the rack, that's integrated with the data center, that's integrated with the power source. And if between each of these boundaries, you have a custom optimized interface that gets you a few percent, Those few percents up and down start adding up, multiplying out, in fact, to something meaningful. And that is exactly what we go after.
AI assessment note: “for us, for let's say our TPUs, we co-design them with a building”
Partly produced feed
D 3 · C 3 · P 3 · Cm 3 3.00
Q of on-site stuff, right? You end up over-provisioning really heavily, and then eventually you get the grid connection, and now what do you do with all this stuff? So Is there, during that bridge power period, are you offering a different level of service somehow, or are you actually provisioning for your two nines, whatever your ultimate reliability requirement is going to be, but from day one with onsite resources?
A It's both. I mean, we basically, I mean, one way to look at it is that, uh, um, most people have trouble, unless they've operated at scale, thinking in terms of these numbers of, like, what's the difference between 99.9 and 99.5 or 99.99 in a given year? And in a given year, they might actually be identical. And so some people are just going to say, I'm gonna roll the dice. I, I hope I get lucky, and sometimes they will, and they actually won't experience any, any issues. Uh, what I would say, though, is that we also look to seeing, okay, beyond some of this bridge power that we're going to need, what are the more permanent sources? Would we use solar, wind, nuclear, other sources that will be permanent, but might not be able to get us all the way to the power capacity that we might need, and then we have to augment with whatever it might, might be turbines, gas, or, or something else.
AI assessment note: “It's both. I mean, we basically, I mean, one way to look at it”