GPU clusters

also referred to as: gpu cluster

7 statements across 5 episodes · 3 bullish · 3 bearish · 5 people on the record · first statement Sep 7, 2023 by Andrew Feldman · across every show →

Everything said about GPU clusters, oldest first

Sep 7, 2023 positive
Assertion Not checkable as stated
Feldman: Cerebras converged a model in 3.5 days after 60-day GPU failure
“We had a situation where they were trying to train on a GPU cluster and they were at 60 days and it wasn't converging and We stood it up, and three and a half days later, their model converged”
Andrew Feldman Sep 7, 2023 ▶ 19:04 No Priors Ep. 31 | With Cerebras CEO Andrew Feldman
Sep 7, 2023 positive
Assertion Not checkable as stated
Feldman: Redistributing AI training takes one keystroke on Cerebras vs GPUs
“In March, we put seven GPT models in the open source community. Everybody else was putting one. Why? Because it's really hard to redistribute work across a GPU cluster. For us, it's one keystroke.”
Andrew Feldman Sep 7, 2023 ▶ 11:13 No Priors Ep. 31 | With Cerebras CEO Andrew Feldman
Mar 21, 2024 neutral
Insight
Srivastava: Inter-rack networking matters less for AI inference than training
“Even the GPU clusters themselves, like, you know, the full training networking is a very, very important Piece to have networking on the racks themselves with inference and matters a little less because you're doing a little bit more on individual GPUs and les…”
Tuhin Srivastava Mar 21, 2024 ▶ 4:57 No Priors Ep 56 | With Baseten CEO and Co-Founder Tuhin Srivastava
May 9, 2024 positive
Insight
Sarah Guo notes overtraining LLMs past optimal compute continues to improve performance
“And if you are meta and you have Somewhere between, you know, 22,000 GPU clusters and 350,000 GPUs available then continuing to train past, like, supposedly optimal points, like, does improve performance apparently and doesn't just fully asymptote as soon as m…”
Sarah Guo May 9, 2024 ▶ 14:25 No Priors Ep. 63 | With Sarah Guo and Elad Gil
Aug 22, 2024 bearish
Prediction Not checkable as stated
Davis: Future AI infrastructure will rely far less on massive GPU clusters
“I think this actually points the way towards like what the infrastructure future might look like. And I think it looks a lot less like everything requiring these big clusters.”
Jared Quincy Davis Aug 22, 2024 ▶ 32:22 No Priors Ep. 77 | With Foundry CEO and Founder Jared Quincy Davis
Aug 22, 2024 negative
Assertion Not checkable as stated
Davis: Large-scale GPU pre-training utilization is sub-80% and often below 50%
“And so even for a lot of the more sophisticated orgs running large pre-trainings at scale, The utilization sub-eighty percent, sometimes less than 50%, actually, depending on how bad of a batch they have and the frequency of failure in the cluster.”
Jared Quincy Davis Aug 22, 2024 ▶ 5:16 No Priors Ep. 77 | With Foundry CEO and Founder Jared Quincy Davis
Aug 29, 2024 bearish
Prediction Not checkable as stated
Garman predicts liquid cooling will make on-prem AI clusters too difficult
“Increasingly, I think that's going to get harder and harder as you move to liquid cooling and larger clusters”
Matt Garman Aug 29, 2024 ▶ 15:19 No Priors Ep. 78 | With AWS CEO Matt Garman
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.