Foundry CEO Jared Quincy Davis discusses how the emergence of compound AI systems and smaller high-quality models reduces the need for ultra-large interconnected training clusters.
Assertion Not checkable as stated
Davis: Large-scale GPU pre-training utilization is sub-80% and often below 50%
“And so even for a lot of the more sophisticated orgs running large pre-trainings at scale,
The utilization sub-eighty percent, sometimes less than 50%, actually, depending on how bad of a batch they have and the frequency of failure in the cluster.”
Opinion
Davis: Current AI cloud is basically co-location, not real cloud
“I think current AI cloud is not cloud in the originally intended sense by any means.
[627] Jared Quincy Davis: so we should pull on that thread, but I'd say right now it's basically co-location.
[631] Jared Quincy Davis: Yeah, it's basically co-location, right…”
Assertion Not checkable as stated
Davis: AI cloud customers are forced into unwanted 3-year GPU contracts
“In AI cloud today, you're kind of forced to get really long-term reservations often three years for a fixed amount of capacity.
[944] Jared Quincy Davis: No one really wants 64 GPUs for three years or a thousand GPUs for three years.
[949] Jared Quincy Davis: …”
Assertion Not checkable as stated
Davis: Public clouds own only basis points of global GPU capacity
“What percentage of the world's GPU petaflock capacity, or exaflock capacity, is kind of owned by the major public clouds. And I've asked this many people, and I typically have gotten guesses, you know, in the high tens of percents, and the only time I got a lo…”
Assertion Not checkable as stated
Davis: NVIDIA H100 GPU utilization is 25% or lower
“By many measures, utilization of these, even H-one hundred systems are kind of state-of-the-art, the most viable, the most precious, et cetera, Is, in my case, it's 20%, 25% or lower, according to some you know, pretty high quality data I've seen from some gre…”
Insight
Davis: AI workloads are shifting from large pre-training to batch inference
“I think people are getting more sophisticated at thinking about cost in a more of a life cycle way, and that's actually leading to the workload shifting from large pre-training more and more towards things like batch inference. Which is actually a really, real…”