Jared Quincy Davis, founder and CEO of Foundry, discusses the geographic and institutional distribution of worldwide GPU compute with Sarah Guo and Elad Gil.
Assertion Not checkable as stated
Davis: Large-scale GPU pre-training utilization is sub-80% and often below 50%
“And so even for a lot of the more sophisticated orgs running large pre-trainings at scale,
The utilization sub-eighty percent, sometimes less than 50%, actually, depending on how bad of a batch they have and the frequency of failure in the cluster.”
Opinion
Davis: Current AI cloud is basically co-location, not real cloud
“I think current AI cloud is not cloud in the originally intended sense by any means.
[627] Jared Quincy Davis: so we should pull on that thread, but I'd say right now it's basically co-location.
[631] Jared Quincy Davis: Yeah, it's basically co-location, right…”
Assertion Not checkable as stated
Davis: AI cloud customers are forced into unwanted 3-year GPU contracts
“In AI cloud today, you're kind of forced to get really long-term reservations often three years for a fixed amount of capacity.
[944] Jared Quincy Davis: No one really wants 64 GPUs for three years or a thousand GPUs for three years.
[949] Jared Quincy Davis: …”
Assertion Not checkable as stated
Davis: NVIDIA H100 GPU utilization is 25% or lower
“By many measures, utilization of these, even H-one hundred systems are kind of state-of-the-art, the most viable, the most precious, et cetera, Is, in my case, it's 20%, 25% or lower, according to some you know, pretty high quality data I've seen from some gre…”
Prediction Not checkable as stated
Davis: Future AI infrastructure will rely far less on massive GPU clusters
“I think this actually points the way towards like what the infrastructure future might look like. And I think it looks a lot less like everything requiring these big clusters.”
Insight
Davis: AI workloads are shifting from large pre-training to batch inference
“I think people are getting more sophisticated at thinking about cost in a more of a life cycle way, and that's actually leading to the workload shifting from large pre-training more and more towards things like batch inference. Which is actually a really, real…”