Assertion Supported AI assessment confidence: 95% certainty 4/5 debate potential 2/5

Davis: Compound AI approach yielded a 3% MMLU performance bump

Jared Quincy Davis · No Priors Ep. 77 | With Foundry CEO and Founder Jared Quincy Davis · Aug 22, 2024 · at 39:55

Foundry CEO Jared Quincy Davis discusses empirical results from his co-authored paper evaluating compound AI systems across benchmark tasks.

0:00 / 0:15exact quote · 16.0s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“The MMLU performance bump was about three percent. And to put that in perspective, the gap between some of the previous best models is often less than one percent. Between, for example, Gemini, 1.5, and Lama 3.1, and things like that. So, actually, 2.8% or three percent is actually a pretty major gap on MMLU.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Jared Quincy Davis

Assertion Not checkable as stated
Davis: Large-scale GPU pre-training utilization is sub-80% and often below 50%
“And so even for a lot of the more sophisticated orgs running large pre-trainings at scale, The utilization sub-eighty percent, sometimes less than 50%, actually, depending on how bad of a batch they have and the frequency of failure in the cluster.”
Jared Quincy Davis Aug 22, 2024 ▶ 5:16 No Priors Ep. 77 | With Foundry CEO and Founder Jared Quincy Davis
Opinion
Davis: Current AI cloud is basically co-location, not real cloud
“I think current AI cloud is not cloud in the originally intended sense by any means. [627] Jared Quincy Davis: so we should pull on that thread, but I'd say right now it's basically co-location. [631] Jared Quincy Davis: Yeah, it's basically co-location, right…”
Jared Quincy Davis Aug 22, 2024 ▶ 10:23 No Priors Ep. 77 | With Foundry CEO and Founder Jared Quincy Davis
Assertion Not checkable as stated
Davis: AI cloud customers are forced into unwanted 3-year GPU contracts
“In AI cloud today, you're kind of forced to get really long-term reservations often three years for a fixed amount of capacity. [944] Jared Quincy Davis: No one really wants 64 GPUs for three years or a thousand GPUs for three years. [949] Jared Quincy Davis: …”
Jared Quincy Davis Aug 22, 2024 ▶ 15:37 No Priors Ep. 77 | With Foundry CEO and Founder Jared Quincy Davis
Assertion Not checkable as stated
Davis: Public clouds own only basis points of global GPU capacity
“What percentage of the world's GPU petaflock capacity, or exaflock capacity, is kind of owned by the major public clouds. And I've asked this many people, and I typically have gotten guesses, you know, in the high tens of percents, and the only time I got a lo…”
Jared Quincy Davis Aug 22, 2024 ▶ 24:27 No Priors Ep. 77 | With Foundry CEO and Founder Jared Quincy Davis
Assertion Not checkable as stated
Davis: NVIDIA H100 GPU utilization is 25% or lower
“By many measures, utilization of these, even H-one hundred systems are kind of state-of-the-art, the most viable, the most precious, et cetera, Is, in my case, it's 20%, 25% or lower, according to some you know, pretty high quality data I've seen from some gre…”
Jared Quincy Davis Aug 22, 2024 ▶ 28:13 No Priors Ep. 77 | With Foundry CEO and Founder Jared Quincy Davis
Prediction Not checkable as stated
Davis: Future AI infrastructure will rely far less on massive GPU clusters
“I think this actually points the way towards like what the infrastructure future might look like. And I think it looks a lot less like everything requiring these big clusters.”
Jared Quincy Davis Aug 22, 2024 ▶ 32:22 No Priors Ep. 77 | With Foundry CEO and Founder Jared Quincy Davis
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.