The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 21 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Assertion Not checkable as stated
Frankle: No Databricks enterprise customer asks for abstract reasoning AI
“I don't think I have a single customer that's asking to, you know, have AI solve abstract reasoning problems.”
Jonathan Frankle Jun 25, 2024 ▶ 1:13:34 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Insight
Frankle: Needle in a Haystack eval fails to measure holistic context usage
“I think the problems with needle in a haystack are well known. You know, it doesn't measure anything real. You're not even testing the model's ability to holistically use the context just to identify one part of the context. So you can do some wacky things to …”
Jonathan Frankle Jun 25, 2024 ▶ 1:17:50 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Assertion Supported
Frankle: Databricks model is uniquely trained purely on Shutterstock data
“So a lot of models have had Shutterstock data incorporated into them, but this is the only model I know of so far where it was, you know, exclusively and specifically trained just on the vanilla Shutterstock data. There was nothing else mixed in. You know, we …”
Jonathan Frankle Jun 25, 2024 ▶ 3:14 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Assertion Supported
Frankle: Dynamic data mixing during pre-training is effective for domain-specific models
“We've had some surprisingly good luck with this. We just released a paper on it. The details matter a lot and it really matters what you're trying to do with the model. But it's been quite effective for us depending on the setting. And certainly when we're thi…”
Jonathan Frankle Jun 25, 2024 ▶ 53:17 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Insight
Frankle: Fault tolerance is missing from fundamental model training primitives
“Fault tolerance is still not really built into any of the fundamental primitives of training models. And so if something breaks, you have to go figure out what broke your job stops. You have to restart your job. It is a nightmare just to get to the point where…”
Jonathan Frankle Jun 25, 2024 ▶ 9:21 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Assertion Not checkable as stated
Frankle: Most AI data centers are retrofitted, not built for high heat
“In data centers that for the most part were not built remotely for this kind of power or heat and have been retrofitted for this. Like failures happen on a good day with normal CPUs. And this is not a good day and not a normal CPU for the most part.”
Jonathan Frankle Jun 25, 2024 ▶ 12:41 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Insight
Frankle: Large-scale model training forces teams to debug the full infrastructure stack
“It's kind of impossible if you're doing training to not go all the way through the entire stack, regardless of what happens. Like somehow I'm still chatting with cloud providers about power contracts, even though the whole point of dealing with the cloud provi…”
Jonathan Frankle Jun 25, 2024 ▶ 23:07 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Insight
Frankle: Training MoE models with FSDP creates severe network bandwidth bottlenecks
“And those models are very demanding when it comes to network bandwidth, at least if you're training them in kind of FSTP zero three style. Where there's just a lot of parameters getting shuffled back and forth and your ratio of kind of compute to amount of dat…”
Jonathan Frankle Jun 25, 2024 ▶ 39:59 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Insight
Frankle: Deep learning log scales can make trends look however you want
“Anything can look however you want it to look if you put it on a log scale to a certain extent. And log, we love our log scales and deep learning for various reasons. Everything looks very clean on a log scale until everything looks very flat on a log scale.”
Jonathan Frankle Jun 25, 2024 ▶ 57:27 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Insight
Frankle: Top AI scientists must tolerate broken infrastructure and imperfect evals
“Like the most successful scientists I see are the ones who are okay operating in a world where everything's going to be broken. And yet we can still cobble things together and make something interesting happen.”
Jonathan Frankle Jun 25, 2024 ▶ 1:07:46 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Assertion Not checkable as stated
Frankle: PDF parsing remains an unsolved problem in 2024
“PDF parsing is still an unsolved problem, even in 20, 24.”
Jonathan Frankle Jun 25, 2024 ▶ 1:21:43 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Assertion Not checkable as stated
Frankle: Text-to-SQL is one of the most impactful LLM use cases for enterprise
“Like it's, you know, text to SQL is still, or like having a model be able to make SQL calls in the backend is actually like one of the single most useful things for my customers. It sounds really boring. Models are really good at it and it moves the needle day…”
Jonathan Frankle Jun 25, 2024 ▶ 1:22:12 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Disclosure
Frankle: Releasing Open-Source Models Is Not Databricks' Core Bread and Butter
“Releasing models open source is not our day-to-day bread and butter. It's kind of a fun reward that we get to do sometimes when we have something really cool to share and a little bit of time and spare GPUs in our hands. But for the most part, everything is go…”
Jonathan Frankle Jun 25, 2024 ▶ 1:28:40 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Disclosure
Frankle: Databricks released text-to-image model with Shutterstock
“Is that we finally released our text image model which has been a year in the making through a collaboration directly with Shutterstock.”
Jonathan Frankle Jun 25, 2024 ▶ 1:32 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Assertion Supported
Frankle: OpenAI, Google, Meta, and Apple have data deals with Shutterstock
“And you know, I, at least I've heard in the news, like opening, I Google, Meta Apple have all called Shutterstock and made those deals.”
Jonathan Frankle Jun 25, 2024 ▶ 3:05 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Assertion Not checkable as stated
Frankle: Porch pirates stole MosaicML InfiniBand cables twice before data center delivery
“Our InfiniBand cables getting stolen from the data center twice, like in boxes before they arrived at the data center, like, you know, porch pirate basically had stolen our InfiniBand cables back when those were hard to come by”
Jonathan Frankle Jun 25, 2024 ▶ 10:54 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Disclosure
Frankle: Databricks runs across six or seven different cloud providers
“Think we're running on like six or seven different clouds right now.”
Jonathan Frankle Jun 25, 2024 ▶ 36:15 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Opinion
Frankle: Networking was the hardest part of training DBRX at scale
“And so actually the networking part of DPRX was the single hardest thing. I think of the entire process, just get MOE training, working at scale across a big cluster.”
Jonathan Frankle Jun 25, 2024 ▶ 40:27 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Assertion Supported
Frankle: Google TPUs provide much higher bandwidth-to-compute ratios
“TPUs have a very different network bandwidth to compute ratio. They have a lot more bandwidth just objectively and TPUs per chip tend to be a little bit less compute intensive and have a little bit less memory.”
Jonathan Frankle Jun 25, 2024 ▶ 41:17 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Prediction Not checkable as stated
Frankle: Multimodal AI models will inevitably require massive context windows
“Once you get into multimodal land, you're just going to end up with giant context. It's kind of unavoidable.”
Jonathan Frankle Jun 25, 2024 ▶ 1:16:55 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Disclosure
Databricks codenamed DBRX Kadabra, teasing a third Alakazam model evolution
“The DBRX small model that we still haven't released yet was called Abra. DBRX was called Kadabra and, you know, there's a third Pokemon in that evolution and that's all I'll say for now.”
Jonathan Frankle Jun 25, 2024 ▶ 1:29:19 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.