Assertion Supported AI assessment confidence: 92% certainty 4/5 debate potential 1/5

Frankle: Google TPUs provide much higher bandwidth-to-compute ratios

Jonathan Frankle · State of the Art: Training 70B LLMs on 10,000 H100 clusters · Jun 25, 2024 · at 41:17

Jonathan Frankle compares Google TPU cluster hardware architecture against typical GPU infrastructure for training foundation models.

0:00 / 0:11exact quote · 11.3s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“TPUs have a very different network bandwidth to compute ratio. They have a lot more bandwidth just objectively and TPUs per chip tend to be a little bit less compute intensive and have a little bit less memory.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Jonathan Frankle

Assertion Not checkable as stated
Frankle: No Databricks enterprise customer asks for abstract reasoning AI
“I don't think I have a single customer that's asking to, you know, have AI solve abstract reasoning problems.”
Jonathan Frankle Jun 25, 2024 ▶ 1:13:34 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Insight
Frankle: Needle in a Haystack eval fails to measure holistic context usage
“I think the problems with needle in a haystack are well known. You know, it doesn't measure anything real. You're not even testing the model's ability to holistically use the context just to identify one part of the context. So you can do some wacky things to …”
Jonathan Frankle Jun 25, 2024 ▶ 1:17:50 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Assertion Supported
Frankle: Databricks model is uniquely trained purely on Shutterstock data
“So a lot of models have had Shutterstock data incorporated into them, but this is the only model I know of so far where it was, you know, exclusively and specifically trained just on the vanilla Shutterstock data. There was nothing else mixed in. You know, we …”
Jonathan Frankle Jun 25, 2024 ▶ 3:14 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Assertion Supported
Frankle: Dynamic data mixing during pre-training is effective for domain-specific models
“We've had some surprisingly good luck with this. We just released a paper on it. The details matter a lot and it really matters what you're trying to do with the model. But it's been quite effective for us depending on the setting. And certainly when we're thi…”
Jonathan Frankle Jun 25, 2024 ▶ 53:17 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Insight
Frankle: Fault tolerance is missing from fundamental model training primitives
“Fault tolerance is still not really built into any of the fundamental primitives of training models. And so if something breaks, you have to go figure out what broke your job stops. You have to restart your job. It is a nightmare just to get to the point where…”
Jonathan Frankle Jun 25, 2024 ▶ 9:21 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Assertion Not checkable as stated
Frankle: Most AI data centers are retrofitted, not built for high heat
“In data centers that for the most part were not built remotely for this kind of power or heat and have been retrofitted for this. Like failures happen on a good day with normal CPUs. And this is not a good day and not a normal CPU for the most part.”
Jonathan Frankle Jun 25, 2024 ▶ 12:41 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.