Databricks

includes Databricks Genie, Databricks AI Runtime, Databricks Marketplace

32 statements across 4 episodes · 19 bullish · 1 bearish · 5 people on the record · first statement Jun 11, 2024 by Mike Conover · said 106 times in 24 episodes since 2023 · across every show →

Mentions by year, the whole family

brought up most by Shawn Wang (36), Matei Zaharia (10), Reynold Xin (7), Jonathan Frankle (7), Ankur Goyal (6), Andy Konwinski (4), Simon Eskildsen (3), Varun Mohan (2)

tap a year for its mentions
0025550102023202420252026episodesmentions
05102023202420252026episodes it came up in
00458102023202420252026episodesmentions per episode
2026 44 mentions in 6 episodes 7 per episode
2025 26 mentions in 9 episodes 3 per episode
2024 35 mentions in 8 episodes 4 per episode
2023 1 mention in 1 episode

every mention, scene by scene, with the transcript →

Everything said about Databricks, oldest first

Jun 11, 2024
Insight
Conover: Repeatable AI annotation pipelines require unsexy people management
“It's one thing to do like a single monolithic push to create a, Training data set like that, or an evaluation corpus, but I think it's another to have a repeatable process, and a lot of that, I think, realistically is pretty unsexy, like, people management wor…”
Mike Conover Jun 11, 2024 ▶ 30:14 How AI is Eating Finance - with Mike Conover of Brightwave
Jun 25, 2024
Insight
Frankle: Large-scale model training forces teams to debug the full infrastructure stack
“It's kind of impossible if you're doing training to not go all the way through the entire stack, regardless of what happens. Like somehow I'm still chatting with cloud providers about power contracts, even though the whole point of dealing with the cloud provi…”
Jonathan Frankle Jun 25, 2024 ▶ 23:07 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Jun 25, 2024 positive
Insight
Frankle: Top AI scientists must tolerate broken infrastructure and imperfect evals
“Like the most successful scientists I see are the ones who are okay operating in a world where everything's going to be broken. And yet we can still cobble things together and make something interesting happen.”
Jonathan Frankle Jun 25, 2024 ▶ 1:07:46 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Jun 25, 2024
Opinion
Frankle: Networking was the hardest part of training DBRX at scale
“And so actually the networking part of DPRX was the single hardest thing. I think of the entire process, just get MOE training, working at scale across a big cluster.”
Jonathan Frankle Jun 25, 2024 ▶ 40:27 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Jun 25, 2024 neutral
Disclosure
Frankle: Databricks runs across six or seven different cloud providers
“Think we're running on like six or seven different clouds right now.”
Jonathan Frankle Jun 25, 2024 ▶ 36:15 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Jun 25, 2024 positive
Assertion Supported
Frankle: Dynamic data mixing during pre-training is effective for domain-specific models
“We've had some surprisingly good luck with this. We just released a paper on it. The details matter a lot and it really matters what you're trying to do with the model. But it's been quite effective for us depending on the setting. And certainly when we're thi…”
Jonathan Frankle Jun 25, 2024 ▶ 53:17 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Jun 25, 2024 neutral
Disclosure
Frankle: Releasing Open-Source Models Is Not Databricks' Core Bread and Butter
“Releasing models open source is not our day-to-day bread and butter. It's kind of a fun reward that we get to do sometimes when we have something really cool to share and a little bit of time and spare GPUs in our hands. But for the most part, everything is go…”
Jonathan Frankle Jun 25, 2024 ▶ 1:28:40 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Jun 25, 2024 positive
Disclosure
Frankle: Databricks released text-to-image model with Shutterstock
“Is that we finally released our text image model which has been a year in the making through a collaboration directly with Shutterstock.”
Jonathan Frankle Jun 25, 2024 ▶ 1:32 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Jun 25, 2024 positive
Assertion Supported
Frankle: Databricks model is uniquely trained purely on Shutterstock data
“So a lot of models have had Shutterstock data incorporated into them, but this is the only model I know of so far where it was, you know, exclusively and specifically trained just on the vanilla Shutterstock data. There was nothing else mixed in. You know, we …”
Jonathan Frankle Jun 25, 2024 ▶ 3:14 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Jun 25, 2024 bullish
Disclosure
Databricks codenamed DBRX Kadabra, teasing a third Alakazam model evolution
“The DBRX small model that we still haven't released yet was called Abra. DBRX was called Kadabra and, you know, there's a third Pokemon in that evolution and that's all I'll say for now.”
Jonathan Frankle Jun 25, 2024 ▶ 1:29:19 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Jun 25, 2024 negative
Assertion Not checkable as stated
Frankle: No Databricks enterprise customer asks for abstract reasoning AI
“I don't think I have a single customer that's asking to, you know, have AI solve abstract reasoning problems.”
Jonathan Frankle Jun 25, 2024 ▶ 1:13:34 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Jun 25, 2024 positive
Assertion Not checkable as stated
Frankle: Text-to-SQL is one of the most impactful LLM use cases for enterprise
“Like it's, you know, text to SQL is still, or like having a model be able to make SQL calls in the backend is actually like one of the single most useful things for my customers. It sounds really boring. Models are really good at it and it moves the needle day…”
Jonathan Frankle Jun 25, 2024 ▶ 1:22:12 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Dec 31, 2025 positive
Assertion Not checkable as stated
Jeff Dean and Databricks founders invested in Konwinski's venture fund
“So the venture funding, there's 50 professors and PhDs, the top names like Jeff Dean and the top faculty at Berkeley and Stanford and my co-founders of Databricks and Perplexity who've invested in the fund.”
Andy Konwinski Dec 31, 2025 ▶ 2:43 [State of Research Funding] Beyond NSF, Slingshots, Open Frontiers — Andy Konwinski, Laude Institute
Dec 31, 2025 positive
Opinion
The National Science Foundation has returned thousands of times its public cost
“It has been probably the best investment the American popular public has ever made. Thousands X return on their investments. I mentioned Google earlier, Databricks, all these researchers came from NSF funding”
Andy Konwinski Dec 31, 2025 ▶ 13:42 [State of Research Funding] Beyond NSF, Slingshots, Open Frontiers — Andy Konwinski, Laude Institute
Jun 24, 2026 bullish
Prediction Not checkable as stated
Zaharia: Open agent hosting layers will win over proprietary alternatives
“Another way to think about it is like, imagine, you know we, our thing wasn't open. We had some kind of agent hosting thing, but it's not open. And then there is an open one. If you're, which one's gonna win in the long run? So like here, because there is this…”
Matei Zaharia Jun 24, 2026 ▶ 12:09 The Agent Cloud: Databricks’ Bet on the Future of AI — Matei Zaharia and Reynold Xin
Jun 24, 2026 positive
Insight
Matei Zaharia: AI agents need stateful contextual security policies, not binary permissions
“A lot of coding agents today have very basic things like you can tell me which tool patterns I'll allow or disallow or whatever. It's like yes or no, but that puts you in a very tough spot. So just as an example, like, should my agent be able to read, you know…”
Matei Zaharia Jun 24, 2026 ▶ 19:46 The Agent Cloud: Databricks’ Bet on the Future of AI — Matei Zaharia and Reynold Xin
Jun 24, 2026 positive
Disclosure
Xin: Databricks uses ML models to optimize database algorithms at runtime
“They use that to build a model. Like a machine learning model, not an L, a machine learning model. Machine learning model basically can very, very quickly tell us how any algorithm and how any implementation will perform for any specific type of queries with v…”
Reynold Xin Jun 24, 2026 ▶ 48:03 The Agent Cloud: Databricks’ Bet on the Future of AI — Matei Zaharia and Reynold Xin
Jun 24, 2026 positive
Assertion Not checkable as stated
Xin: Databricks launches 15 to 16 million VMs per day across clouds
“So we, on the analytics side, I think we launched Maybe 50 or sixteen million virtual machines a day across all three clouds, so which are one of the biggest compute orchestrators out there. Stuff for sure for CPU compute.”
Reynold Xin Jun 24, 2026 ▶ 18:16 The Agent Cloud: Databricks’ Bet on the Future of AI — Matei Zaharia and Reynold Xin
Jun 24, 2026 positive
Insight
Zaharia: Scaling down from bulk ingestion to serving is easier than scaling up
“It turned out that, you know, it's easier to go from that bod thing that's really good at the scale and ingesting and super low cost and create versions in it that have the speed and features of the, you know, super easy to use, like smaller data for business …”
Matei Zaharia Jun 24, 2026 ▶ 55:55 The Agent Cloud: Databricks’ Bet on the Future of AI — Matei Zaharia and Reynold Xin
Jun 24, 2026 positive
Assertion Supported
Xin: Transcoding database rows to Parquet speeds object storage writes with zero compromise
“And as a matter of fact, once you transcode the data compresses better. So from those services writing to, for example, S three or other data lake, like object stores, you can actually write them faster because now they are now smaller. So there's no. Overhead…”
Reynold Xin Jun 24, 2026 ▶ 36:50 The Agent Cloud: Databricks’ Bet on the Future of AI — Matei Zaharia and Reynold Xin
Jun 24, 2026
Insight
Xin: Optimizing Solely for Tech Companies Impedes Traditional Enterprise Scaling
“One of the challenges I think we probably see, and maybe many newer generation companies are seeing is, so tech companies are very, very different from non-tech companies or traditional enterprises. And if you optimize everything just for tech companies, you m…”
Reynold Xin Jun 24, 2026 ▶ 42:25 The Agent Cloud: Databricks’ Bet on the Future of AI — Matei Zaharia and Reynold Xin
Jun 24, 2026 positive
Insight
Zaharia: Software with integration network effects should be open source
“One, so, I mean, one of the reasons to open source something is if you think it's a layer that will actually, there'll be some network effect. It'll benefit from many people collaborating on it.”
Matei Zaharia Jun 24, 2026 ▶ 11:20 The Agent Cloud: Databricks’ Bet on the Future of AI — Matei Zaharia and Reynold Xin
Jun 24, 2026 positive
Insight
Xin: Unifying storage delivers 99% of HTAP database benefits
“HTAP wants to build a single engine for both. We think you can get 99% of what you need by unifying the storage and just have a single storage layer.”
Reynold Xin Jun 24, 2026 ▶ 33:16 The Agent Cloud: Databricks’ Bet on the Future of AI — Matei Zaharia and Reynold Xin
Jun 24, 2026 positive
Disclosure
Zaharia: Databricks built Isaac as an internal wrapper for Claude Code and Codex
“We have a really great dev info team. They built something called Isaac that's basically like a wrapper on cloud code and codex and let's you use them either on the web and like, Sandboxes or just on your dev machine or on your laptop or whatever.”
Matei Zaharia Jun 24, 2026 ▶ 3:47 The Agent Cloud: Databricks’ Bet on the Future of AI — Matei Zaharia and Reynold Xin
Jun 24, 2026 neutral
Assertion Not checkable as stated
Xin: Databricks teams internally built five or six redundant agent frameworks
“I think we had like five or six different agentic frameworks built by every different team. They do all do more or less the same thing.”
Reynold Xin Jun 24, 2026 ▶ 10:59 The Agent Cloud: Databricks’ Bet on the Future of AI — Matei Zaharia and Reynold Xin
Jun 24, 2026 neutral
Disclosure
Zaharia: Databricks gives internal developers unlimited AI token spend
“It's unlimited, but we do you know, we use our own product to like analyze the traces and stuff, and we have a team that's, you know, looking to optimize and to see if anyone's doing something weird.”
Matei Zaharia Jun 24, 2026 ▶ 24:22 The Agent Cloud: Databricks’ Bet on the Future of AI — Matei Zaharia and Reynold Xin
Jun 24, 2026
Insight
Xin: Overfitting to Initial Customers Has Far Lower Downside Than Boiling the Ocean
“I think the industry has a sense of, hey, maybe if you overfit to like one or two customers, it's going to be really bad for you. But I think the downside overfitting is much smaller than the upside itself. And if you sort of try to be too ambitious and boil t…”
Reynold Xin Jun 24, 2026 ▶ 41:55 The Agent Cloud: Databricks’ Bet on the Future of AI — Matei Zaharia and Reynold Xin
Jun 24, 2026 neutral
Insight
Xin: Customer friction with vendor sprawl forced Modern Data Stack consolidation
“What people eventually run into, It's kind of a question of a unification and consolidation is, hey, do you really need to chop all of this into different pieces and work with so many different vendors and platforms in order to get like a very simple visualiza…”
Reynold Xin Jun 24, 2026 ▶ 14:40 The Agent Cloud: Databricks’ Bet on the Future of AI — Matei Zaharia and Reynold Xin
Jun 24, 2026 positive
Disclosure
Zaharia: Omnigent was built to allow teams to securely share custom agent setups
“We had a lot of engineers building, you know, their own Vibe coding setup. But then the other thing they all said is like, Hey, I built something that's amazing for me. But like no one else on the team can use it because I don't have a server to collaborate. A…”
Matei Zaharia Jun 24, 2026 ▶ 9:12 The Agent Cloud: Databricks’ Bet on the Future of AI — Matei Zaharia and Reynold Xin
Jun 24, 2026 bullish
Assertion Not checkable as stated
Zaharia: Databricks' document vision model is ~100x cheaper than frontier models
“Our team built this document sort of vision model that takes a page and gives you back a nice JSON with all the components. And it's very competitive. It's like probably like a hundred X cheaper than those frontier models and still better.”
Matei Zaharia Jun 24, 2026 ▶ 1:01:56 The Agent Cloud: Databricks’ Bet on the Future of AI — Matei Zaharia and Reynold Xin
Jun 24, 2026 neutral
Assertion Supported
Xin: Every major analytics database engine in traction is a decade old
“Actually, every single database engine out there, especially on the analytics side, are kind of a decade old. Pretty much everything that had reasonable traction are about a decade old.”
Reynold Xin Jun 24, 2026 ▶ 45:08 The Agent Cloud: Databricks’ Bet on the Future of AI — Matei Zaharia and Reynold Xin
Jun 24, 2026 neutral
Disclosure
Zaharia: Databricks abandons frontier AI models to focus on agent systems
“Even though we did launch open source model DBRX, and, you know, we went up to, like, sort of above the LAMA-R III scale, we decided that we really want to focus on, there'll be so many people releasing models, and instead of doing the general model where, lik…”
Matei Zaharia Jun 24, 2026 ▶ 1:00:03 The Agent Cloud: Databricks’ Bet on the Future of AI — Matei Zaharia and Reynold Xin
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.