“TPUs have a very different network bandwidth to compute ratio. They have a lot more bandwidth just objectively and TPUs per chip tend to be a little bit less compute intensive and have a little bit less memory.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Jonathan Frankle
AssertionNot checkable as stated
Frankle: No Databricks enterprise customer asks for abstract reasoning AI
“I don't think I have a single customer that's asking to, you know, have AI solve abstract reasoning problems.”
Jonathan FrankleJun 25, 2024▶ 1:13:34State of the Art: Training 70B LLMs on 10,000 H100 clusters
Insight
Frankle: Needle in a Haystack eval fails to measure holistic context usage
“I think the problems with needle in a haystack are well known. You know, it doesn't measure anything real. You're not even testing the model's ability to holistically use the context just to identify one part of the context. So you can do some wacky things to …”
Jonathan FrankleJun 25, 2024▶ 1:17:50State of the Art: Training 70B LLMs on 10,000 H100 clusters
AssertionSupported
Frankle: Databricks model is uniquely trained purely on Shutterstock data
“So a lot of models have had Shutterstock data incorporated into them, but this is the only model I know of so far where it was, you know, exclusively and specifically trained just on the vanilla Shutterstock data. There was nothing else mixed in. You know, we …”
Jonathan FrankleJun 25, 2024▶ 3:14State of the Art: Training 70B LLMs on 10,000 H100 clusters
AssertionSupported
Frankle: Dynamic data mixing during pre-training is effective for domain-specific models
“We've had some surprisingly good luck with this. We just released a paper on it. The details matter a lot and it really matters what you're trying to do with the model. But it's been quite effective for us depending on the setting. And certainly when we're thi…”
Jonathan FrankleJun 25, 2024▶ 53:17State of the Art: Training 70B LLMs on 10,000 H100 clusters
Insight
Frankle: Fault tolerance is missing from fundamental model training primitives
“Fault tolerance is still not really built into any of the fundamental primitives of training models. And so if something breaks, you have to go figure out what broke your job stops. You have to restart your job. It is a nightmare just to get to the point where…”
Jonathan FrankleJun 25, 2024▶ 9:21State of the Art: Training 70B LLMs on 10,000 H100 clusters
AssertionNot checkable as stated
Frankle: Most AI data centers are retrofitted, not built for high heat
“In data centers that for the most part were not built remotely for this kind of power or heat and have been retrofitted for this. Like failures happen on a good day with normal CPUs. And this is not a good day and not a normal CPU for the most part.”
Jonathan FrankleJun 25, 2024▶ 12:41State of the Art: Training 70B LLMs on 10,000 H100 clusters
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 200 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.