Jonathan Frankle

Chief AI Scientist, Databricks · 1 appearance on the record.

computed by AI from the episodes · how this works → · full disclaimer →

Jonathan Frankle serves as the Chief AI Scientist at Databricks. He focuses on training large-scale AI models on massive computing clusters.

21statements → 10claims → 4claims resolved → 100%fully supported → 3.67/5average certainty → 1.71/5average debate potential → 14said about them ↓

4 supported 0 partly supported 0 contradicted 6 not checkable as stated how the 10 claims stand · each chip opens the sources

1 prediction · 9 assertions · 1 opinion · 6 insights · 4 disclosures · every statement was checked. The prediction and assertions are the 10 claims: statements the public record can support or contradict. 4 are resolved, and 6 name no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Jonathan argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
Frankle: Databricks model is uniquely trained purely on Shutterstock data
“So a lot of models have had Shutterstock data incorporated into them, but this is the only model I know of so far where it was, you know, exclusively and specifically trained just on the vanilla Shutterstock data. There was nothing else mixed in. You know, we …”
Jonathan Frankle Jun 25, 2024 ▶ 3:14 State of the Art: Training 70B LLMs on 10,000 H100 clusters

Expressed certainty vs assessment result

none yet certainty 1
100% certainty 2
100% certainty 3
100% certainty 4
none yet certainty 5

weighted support: a fully supported claim counts one, a partly supported claim counts half. Each filled bar is clickable and opens exactly those claims; "none yet" means nothing said at that certainty level has resolved yet

Everything Jonathan Frankle said on Latent Space that made the record, most notable first. Filter by type, assessment or year in the ledger →

Assertion Not checkable as stated
Frankle: No Databricks enterprise customer asks for abstract reasoning AI
“I don't think I have a single customer that's asking to, you know, have AI solve abstract reasoning problems.”
Jonathan Frankle Jun 25, 2024 ▶ 1:13:34 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Insight
Frankle: Needle in a Haystack eval fails to measure holistic context usage
“I think the problems with needle in a haystack are well known. You know, it doesn't measure anything real. You're not even testing the model's ability to holistically use the context just to identify one part of the context. So you can do some wacky things to …”
Jonathan Frankle Jun 25, 2024 ▶ 1:17:50 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Assertion Supported
Frankle: Databricks model is uniquely trained purely on Shutterstock data
“So a lot of models have had Shutterstock data incorporated into them, but this is the only model I know of so far where it was, you know, exclusively and specifically trained just on the vanilla Shutterstock data. There was nothing else mixed in. You know, we …”
Jonathan Frankle Jun 25, 2024 ▶ 3:14 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Assertion Supported
Frankle: Dynamic data mixing during pre-training is effective for domain-specific models
“We've had some surprisingly good luck with this. We just released a paper on it. The details matter a lot and it really matters what you're trying to do with the model. But it's been quite effective for us depending on the setting. And certainly when we're thi…”
Jonathan Frankle Jun 25, 2024 ▶ 53:17 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Insight
Frankle: Fault tolerance is missing from fundamental model training primitives
“Fault tolerance is still not really built into any of the fundamental primitives of training models. And so if something breaks, you have to go figure out what broke your job stops. You have to restart your job. It is a nightmare just to get to the point where…”
Jonathan Frankle Jun 25, 2024 ▶ 9:21 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Assertion Not checkable as stated
Frankle: Most AI data centers are retrofitted, not built for high heat
“In data centers that for the most part were not built remotely for this kind of power or heat and have been retrofitted for this. Like failures happen on a good day with normal CPUs. And this is not a good day and not a normal CPU for the most part.”
Jonathan Frankle Jun 25, 2024 ▶ 12:41 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Insight
Frankle: Large-scale model training forces teams to debug the full infrastructure stack
“It's kind of impossible if you're doing training to not go all the way through the entire stack, regardless of what happens. Like somehow I'm still chatting with cloud providers about power contracts, even though the whole point of dealing with the cloud provi…”
Jonathan Frankle Jun 25, 2024 ▶ 23:07 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Insight
Frankle: Training MoE models with FSDP creates severe network bandwidth bottlenecks
“And those models are very demanding when it comes to network bandwidth, at least if you're training them in kind of FSTP zero three style. Where there's just a lot of parameters getting shuffled back and forth and your ratio of kind of compute to amount of dat…”
Jonathan Frankle Jun 25, 2024 ▶ 39:59 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Insight
Frankle: Deep learning log scales can make trends look however you want
“Anything can look however you want it to look if you put it on a log scale to a certain extent. And log, we love our log scales and deep learning for various reasons. Everything looks very clean on a log scale until everything looks very flat on a log scale.”
Jonathan Frankle Jun 25, 2024 ▶ 57:27 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Insight
Frankle: Top AI scientists must tolerate broken infrastructure and imperfect evals
“Like the most successful scientists I see are the ones who are okay operating in a world where everything's going to be broken. And yet we can still cobble things together and make something interesting happen.”
Jonathan Frankle Jun 25, 2024 ▶ 1:07:46 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Assertion Not checkable as stated
Frankle: PDF parsing remains an unsolved problem in 2024
“PDF parsing is still an unsolved problem, even in 20, 24.”
Jonathan Frankle Jun 25, 2024 ▶ 1:21:43 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Assertion Not checkable as stated
Frankle: Text-to-SQL is one of the most impactful LLM use cases for enterprise
“Like it's, you know, text to SQL is still, or like having a model be able to make SQL calls in the backend is actually like one of the single most useful things for my customers. It sounds really boring. Models are really good at it and it moves the needle day…”
Jonathan Frankle Jun 25, 2024 ▶ 1:22:12 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Disclosure
Frankle: Releasing Open-Source Models Is Not Databricks' Core Bread and Butter
“Releasing models open source is not our day-to-day bread and butter. It's kind of a fun reward that we get to do sometimes when we have something really cool to share and a little bit of time and spare GPUs in our hands. But for the most part, everything is go…”
Jonathan Frankle Jun 25, 2024 ▶ 1:28:40 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Disclosure
Frankle: Databricks released text-to-image model with Shutterstock
“Is that we finally released our text image model which has been a year in the making through a collaboration directly with Shutterstock.”
Jonathan Frankle Jun 25, 2024 ▶ 1:32 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Assertion Supported
Frankle: OpenAI, Google, Meta, and Apple have data deals with Shutterstock
“And you know, I, at least I've heard in the news, like opening, I Google, Meta Apple have all called Shutterstock and made those deals.”
Jonathan Frankle Jun 25, 2024 ▶ 3:05 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Assertion Not checkable as stated
Frankle: Porch pirates stole MosaicML InfiniBand cables twice before data center delivery
“Our InfiniBand cables getting stolen from the data center twice, like in boxes before they arrived at the data center, like, you know, porch pirate basically had stolen our InfiniBand cables back when those were hard to come by”
Jonathan Frankle Jun 25, 2024 ▶ 10:54 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Disclosure
Frankle: Databricks runs across six or seven different cloud providers
“Think we're running on like six or seven different clouds right now.”
Jonathan Frankle Jun 25, 2024 ▶ 36:15 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Opinion
Frankle: Networking was the hardest part of training DBRX at scale
“And so actually the networking part of DPRX was the single hardest thing. I think of the entire process, just get MOE training, working at scale across a big cluster.”
Jonathan Frankle Jun 25, 2024 ▶ 40:27 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Assertion Supported
Frankle: Google TPUs provide much higher bandwidth-to-compute ratios
“TPUs have a very different network bandwidth to compute ratio. They have a lot more bandwidth just objectively and TPUs per chip tend to be a little bit less compute intensive and have a little bit less memory.”
Jonathan Frankle Jun 25, 2024 ▶ 41:17 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Prediction Not checkable as stated
Frankle: Multimodal AI models will inevitably require massive context windows
“Once you get into multimodal land, you're just going to end up with giant context. It's kind of unavoidable.”
Jonathan Frankle Jun 25, 2024 ▶ 1:16:55 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Disclosure
Databricks codenamed DBRX Kadabra, teasing a third Alakazam model evolution
“The DBRX small model that we still haven't released yet was called Abra. DBRX was called Kadabra and, you know, there's a third Pokemon in that evolution and that's all I'll say for now.”
Jonathan Frankle Jun 25, 2024 ▶ 1:29:19 State of the Art: Training 70B LLMs on 10,000 H100 clusters

The other half of the tape: Jonathan Frankle's own voice is left out of every number here. Other people bring the name up 13 times in 3 episodes on Latent Space. 1 statement on the record names them. every mention, with the transcript →

Who brings them up most Alessio Fanelli 5Shawn Wang 4

Statements about Jonathan Frankle, by other people (1)

Insight
Morcos: Lottery ticket initializations fail because they are data dependent
“We actually found out that the problem was that the lottery ticket was actually data dependent. And that was where the fundamental problem came. That as soon as you change the data distribution a little bit, like the winning tickets changed in a really big way…”
Ari Morcos Aug 29, 2025 ▶ 1:01:14 Better Data is All You Need — Ari Morcos, Datology

Every mention by year

tap a year for its mentions
0051102202320242025episodesmentions
012202320242025episodes it came up in
002.5152202320242025episodesmentions per episode
2025 9 mentions in 2 episodes 5 per episode
2023 4 mentions in 1 episode

Appearances (1)

EpisodeDateSpeaking time
State of the Art: Training 70B LLMs on 10,000 H100 clusters Jun 25, 2024 27m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.