Foundry CEO Jared Quincy Davis discusses the true utilization rates of GPU clusters during end-to-end foundation model pre-training.
Opinion
Davis: Current AI cloud is basically co-location, not real cloud
“I think current AI cloud is not cloud in the originally intended sense by any means.
[627] Jared Quincy Davis: so we should pull on that thread, but I'd say right now it's basically co-location.
[631] Jared Quincy Davis: Yeah, it's basically co-location, right…”
Assertion Not checkable as stated
Davis: AI cloud customers are forced into unwanted 3-year GPU contracts
“In AI cloud today, you're kind of forced to get really long-term reservations often three years for a fixed amount of capacity.
[944] Jared Quincy Davis: No one really wants 64 GPUs for three years or a thousand GPUs for three years.
[949] Jared Quincy Davis: …”
Assertion Not checkable as stated
Davis: Public clouds own only basis points of global GPU capacity
“What percentage of the world's GPU petaflock capacity, or exaflock capacity, is kind of owned by the major public clouds. And I've asked this many people, and I typically have gotten guesses, you know, in the high tens of percents, and the only time I got a lo…”
Assertion Not checkable as stated
Davis: NVIDIA H100 GPU utilization is 25% or lower
“By many measures, utilization of these, even H-one hundred systems are kind of state-of-the-art, the most viable, the most precious, et cetera, Is, in my case, it's 20%, 25% or lower, according to some you know, pretty high quality data I've seen from some gre…”
Prediction Not checkable as stated
Davis: Future AI infrastructure will rely far less on massive GPU clusters
“I think this actually points the way towards like what the infrastructure future might look like. And I think it looks a lot less like everything requiring these big clusters.”
Insight
Davis: AI workloads are shifting from large pre-training to batch inference
“I think people are getting more sophisticated at thinking about cost in a more of a life cycle way, and that's actually leading to the workload shifting from large pre-training more and more towards things like batch inference. Which is actually a really, real…”