Assertion Not checkable as stated
Davis: Large-scale GPU pre-training utilization is sub-80% and often below 50%
“And so even for a lot of the more sophisticated orgs running large pre-trainings at scale,
The utilization sub-eighty percent, sometimes less than 50%, actually, depending on how bad of a batch they have and the frequency of failure in the cluster.”
Opinion
Davis: Current AI cloud is basically co-location, not real cloud
“I think current AI cloud is not cloud in the originally intended sense by any means.
[627] Jared Quincy Davis: so we should pull on that thread, but I'd say right now it's basically co-location.
[631] Jared Quincy Davis: Yeah, it's basically co-location, right…”
Assertion Not checkable as stated
Davis: AI cloud customers are forced into unwanted 3-year GPU contracts
“In AI cloud today, you're kind of forced to get really long-term reservations often three years for a fixed amount of capacity.
[944] Jared Quincy Davis: No one really wants 64 GPUs for three years or a thousand GPUs for three years.
[949] Jared Quincy Davis: …”
Assertion Not checkable as stated
Davis: Public clouds own only basis points of global GPU capacity
“What percentage of the world's GPU petaflock capacity, or exaflock capacity, is kind of owned by the major public clouds. And I've asked this many people, and I typically have gotten guesses, you know, in the high tens of percents, and the only time I got a lo…”
Assertion Not checkable as stated
Davis: NVIDIA H100 GPU utilization is 25% or lower
“By many measures, utilization of these, even H-one hundred systems are kind of state-of-the-art, the most viable, the most precious, et cetera, Is, in my case, it's 20%, 25% or lower, according to some you know, pretty high quality data I've seen from some gre…”
Prediction Not checkable as stated
Davis: Future AI infrastructure will rely far less on massive GPU clusters
“I think this actually points the way towards like what the infrastructure future might look like. And I think it looks a lot less like everything requiring these big clusters.”
Insight
Davis: AI workloads are shifting from large pre-training to batch inference
“I think people are getting more sophisticated at thinking about cost in a more of a life cycle way, and that's actually leading to the workload shifting from large pre-training more and more towards things like batch inference. Which is actually a really, real…”
Assertion Partly supported
Davis: OpenAI Had 400 Employees and $13B of Compute When Launching ChatGPT
“In OpenAI's case, you know, there were only 400 people, but had thirteen billion dollars worth of compute, you know, which is quite a bit of computational scale there.”
Assertion Not checkable as stated
Davis: AI teams hold 10% to 20% of GPUs idle as healing buffer
“And so one of the consequences of that is that it's very common now to hold aside 10 to 20% minimum of the GPUs that a team has as buffer, as healing buffer, in case of a failure so you can slot something else in to keep the training workload running, right?”
Assertion Supported
Davis: Peak Ethereum mining equaled 10 to 20 million V100 GPUs
“The very peak, the tippy top of Ethereum. How many V-One hundred equivalents were there, given there were 10,000 for two weeks for GPT-III? ... It was about 10 to twenty million.”
Assertion Partly supported
Davis: iPhone 15 Pro has more FP16 FLOPS than an NVIDIA V100
“Actually, an iPhone 15 pro now is actually stronger than a V-one hundred, as a funny example. It has about 35 teraflops in every 16, I believe where a V-one hundred is around 30”
Assertion Supported
Davis: Compound AI approach yielded a 3% MMLU performance bump
“The MMLU performance bump was about three percent. And to put that in perspective, the gap between some of the previous best models is often less than one percent. Between, for example, Gemini, 1.5, and Lama 3.1, and things like that. So, actually, 2.8% or thr…”
Prediction Not checkable as stated
Davis: Democratizing AI Compute Will Increase AlphaFold-Level Breakthroughs 10x to 100x
“I think it would increase the frequency of events like AlphaFold II by 10 X, a hundred X, or maybe even more super linearly.”
Assertion Supported
Davis: Probability of a GPU supercomputer running weeks without failure is zero
“And so because you have millions, perhaps, of individual components in this supercomputer, the probability that it will run for weeks on end, and this is basically a verbatim quote from Jensen's keynote is basically zero.”
Opinion
Davis: NVIDIA's Mellanox buyout is among the best acquisitions in history
“And their acquisition of Mellanox was one of the better of all time, arguably, from a market cap creation perspective.”
Assertion Supported
Davis: AI compute market lacks futures and hedging mechanisms found in commodities
“The markets aren't mature enough that there's any analog. So what we have in other domains, like in commodities markets like wheat, oil, et cetera, where you can buy options and futures and hedge and sell back and things like that. It's kind of a pretty, Still…”
Assertion Supported
Davis: Compound AI approach increased prime factorization accuracy from 3.7% to 36.6%
“And so, you know, we did kind of some preliminary investigations here, and we were able to, in one case, a prime factorization, you know, kind of 10 x the performance, go from 3.7% to 36.6%. On prime factorization, which is pretty hard, kind of factor, you kno…”
Prediction Not checkable as stated
Davis: AI systems will make millions of heterogeneous model calls per question
“I think that what we'll see people doing is kind of composing, this sounds funny, but massive networks, Where maybe each stage in the network will basically be maybe some best of A, best of K component with many, many calls to different language models, you kn…”
Assertion Partly supported
Davis: DeepMind built AlphaFold 2 with a team of just 18 people
“I think that one of the things that was so remarkable to me about AlphaFold II is, initially it was a really small team, you know, three, and then later, 18 people or so, and they solved what was kind of a fifty-year grand challenge in biology, which is a pret…”
Assertion Supported
Davis: OpenAI trained GPT-3 on 10,000 V100s for 14.6 days
“GPT-III was trained on 10,000 V-one-hundred GPUs in an interconnected cluster in Azure for about 14.6 days.”