GPU Clusters
topic on 9 shows · 23 statements across 19 episodes
Latent Space
No Priors
Invest Like the Best
Catalyst
the MAD Podcast
the a16z Podcast
All-In
TBPN
20VC
23 statements about GPU Clusters, every show
Global GPU Utilization Is Actually Far Worse Than xAI's Clusters
“We all make fun of XAI for having, you know, some challenges with total flop utilization on its clusters, but the reality for the rest of the world is it's far worse. A ton of GPUs just sit in warehouses or sit In private pools allocated to a specific customer…”
Intrator: Compute decommoditizes at scale as clusters grow for frontier AI models
“Computing decommoditizes its scale, right? Like, when, you know, anybody can run a GPU, but can you run a cluster that's large enough to train a model that can change the world?”
Lockmiller: Syncing massive AI GPU clusters causes extreme power draw fluctuations
“When you're deploying giant clusters of GPUs you know, the entire data center is acting as a single computer. It's running one single workload that's training some breakthrough foundational model. And what that results in is actually massive load fluctuations …”
AI training compute cycles cause massive power fluctuations unlike inference workloads
“And some of the people in the energy space may know that there are massive energy fluctuations or power fluctuations we will see in data center usage when the GPUs go from this computational intensive phase where you're learning the model weights to this commu…”
Patel: Approximately ten 100,000-GPU AI clusters exist globally
“There are like 10 hundred K GPU clusters in the world.”
Ginn: US needs the world's largest GPU cluster as a deterrent
“I think America should have the largest GPU cluster in the world as a deterrent.”
Conrad: Spot GPU cluster utilization nears 100% through dynamic price clearing
“Assuming there are not, like, hardware problems or software problems, the utilization rate is, like, near a hundred percent, because the price dips until the utilization is a hundred percent.”
Baker: CoreWeave runs large GPU clusters as well as anyone
“And I do think core weave runs these big GPU clusters as well as anyone. And there aren't that many people on planet earth who can run them well.”
Captions parallelizes video rendering via 25-frame segments and overlapping GPU clusters
“We were splitting the video into like 25 frame segments basically, and 25 is 25 FPS. So that's basically a second. And then we were generating overlaps on both sides, four frame overlaps with the next segment. And then we would check for differences, the overl…”
Coogan: Next-gen AI training requires bespoke hundred-thousand-GPU data centers
“I think we're in, like, the custom data center era, where, like, it has to be bespoke, it has to be hundreds of thousands of GPUs.”
Crusoe is building large GPU clusters in Iceland using clean power
“We're doing quite a bit in Iceland where geothermal and hydropower is low cost, clean, and abundant. And we're able to sort of develop these large GPU clusters in Iceland powered entirely by clean energy.”
Lockmiller: Synchronized GPU clusters create extreme power spikes like breathing
“The cluster itself is like moving in, in, in sync with one another. So, you know, you're sort of, ah, all of the data is sort of being broadcast out across this high performance network, you know, to the GPUs. They're running their compute workloads and then p…”
Garman predicts liquid cooling will make on-prem AI clusters too difficult
“Increasingly, I think that's going to get harder and harder as you move to liquid cooling and larger clusters”
Davis: Large-scale GPU pre-training utilization is sub-80% and often below 50%
“And so even for a lot of the more sophisticated orgs running large pre-trainings at scale,
The utilization sub-eighty percent, sometimes less than 50%, actually, depending on how bad of a batch they have and the frequency of failure in the cluster.”
Davis: Future AI infrastructure will rely far less on massive GPU clusters
“I think this actually points the way towards like what the infrastructure future might look like. And I think it looks a lot less like everything requiring these big clusters.”
Albrecht: Imbue's GPU cluster failure rate is well below industry 3% benchmark
“The number that we've heard from other people is like they're having about three percent. I don't think we're experiencing failure rates that are that high. I think ours is actually quite a bit lower than that, probably because we've taken the time to like dig…”
Albrecht: 4,000-GPU Three-Tier Cluster Requires 12,000 Cables and 24,000 Plugs
“Like to bring up this cluster you know, with 4000 GPUs and three tier networking, networking architecture, you have 12,000 cables.
So that's 24,000 things that need to be plugged in.”
Srinivas: Foundation model training requires three-year advance GPU commitments
“Because the way it works is you have to pay three years in advance to get a big cluster. Like you have to commit to that. It's not like all the money goes away immediately, but you have to commit to three year To get like, you know, thousands of GPUs at once i…”
Sarah Guo notes overtraining LLMs past optimal compute continues to improve performance
“And if you are meta and you have Somewhere between, you know, 22,000 GPU clusters and 350,000 GPUs available then continuing to train past, like, supposedly optimal points, like, does improve performance apparently and doesn't just fully asymptote as soon as m…”
Srivastava: Inter-rack networking matters less for AI inference than training
“Even the GPU clusters themselves, like, you know, the full training networking is a very, very important Piece to have networking on the racks themselves with inference and matters a little less because you're doing a little bit more on individual GPUs and les…”
Buying Dedicated GPU Clusters Offers Better Economics for Model Training
“In training, I think, you know, there's less software differentiation. So in training, I think there's certainly, like, better economics of, like, buying big clusters.”