Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 5 · Cm 5 5.00
Q anyone that may be curious about the world of data infrastructure, and then we can go into all sorts of, uh, technical details, but like to start with, what world do you operate in if, if we think of all of this as, as, uh, you know, databases, so databases, data warehouses, how would you sort of compare and contrast the, the, the, the various, uh, databases of the world?
A So I think of the world, uh, the database world As being really divided into two halves, uh, the analytical side and the transactional side. The analytical side are systems that are built to be very read optimized, so reading data as fast as possible. The transactional world being more oriented towards writing data, uh, fast and consistently. So when you think of a transactional system that's maybe powering an application, you know, a, a, a canonical example would be like a, an ATM that needs to record a debit and a credit Uh, very quickly and has to be consistent every single time. An analytical application would be something like, uh, how many customers bought product X last, last year? And slicing and dicing the demographic profile of your customers and understanding their journey, you know, through the website all the way to a transaction at the end. Um, and so those are, those are broadly speaking the, the, the two worlds.
AI assessment note: “database world As being really divided into two halves, uh, the analytical side and the transactional side.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q And like, how do you, how do you, yeah, how do you position?
A Yeah, I think it necessitates being, uh, even crisper on the differentiation between those things, and so for us, there, there are a few things that we point to with customers. Number one, we're one of the only hybrid players, so Databricks and Snowflake are cloud only, so if you happen to have data on-prem, uh, we're pretty much your only bet, and you know, it just so happens that that turns out to be most of the Fortune 500, uh, almost the entirety of the financial Social services sector in particular, and so we do a lot of business in those industries as a result. Um, the second thing that helps differentiate us is the openness of the platform. So we're an open engine querying open formats. Uh, and while there's been widespread embrace, I would say, especially last year in 2024, around open formats and Iceberg in particular really winning that format war, um, that's new. That's new for this industry. We've been doing this Forever, though, and that's, you know, the first queries run on Iceberg were, were Trino or Presto queries. So, um, that pairing in the open source community of Trino and Iceberg, uh, has been kind of a reference architecture for years at this point, and that gives us an advantage because we can help manage your Iceberg deployment holistically. Everything from streaming, you know, ingest, uh, uh, loading that data into Iceberg tables, Maintaining it, doing …
AI assessment note: “for us, there, there are a few things that we point to with customers”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q it's been fascinating actually to see the, the Dell like stock price. Yes. Everybody's talked about Nvidia. Yeah. You know, for all the right reasons, but, um, it's like amazing how Dell has like the quintessential on-prem player, uh, has, uh, seen their fortunes accelerate as well last year. Um, okay, great. So that's, um, Starburst Enterprise and, um, Starburst Galaxy, the managed version had a How does that work?
A Yeah, so that's a classic, you know, SAS, um, uh, product that's hosted and managed by us. It is connecting to your storage, uh, so it's your own S three buckets, your own, you know, RDS, your own MySQL database, uh, but the compute and the control plane is managed by us, and so we're able to offer a very seamless, easy to use, turnkey, push button type of approach, while still giving you all the performance and functionality Uh, that you need and that you're looking for, and that product has actually evolved very, very quickly for us. Um, we've been able to get a lot of new interesting features and functionality there, and one of the cases, one of the use cases where we've seen a lot of, uh, adoption and interest is actually customers who are building their own data applications and using, uh, Galaxy as the embedded engine, uh, where, you know, there's, uh, Data analytics portion of the SaaS app that they provide to their customers and the analytics are essentially powered by our engine. And, you know, we were digging into like, you know, why are they choosing us, uh, for this? And I think it goes down to, you know, if you're gonna be part of, uh, someone else's margin, essentially their cogs, right? Uh, you need good TCO. I think what we're seeing is, um, data application developers are choosing iceberg for storage because that gives them a lot of flexibility. They want Uh, s…
AI assessment note: “that's a classic, you know, SAS, um, uh, product that's hosted and managed by us.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q maybe we'll, we'll, we'll put in the video, uh, where, you know, you have the classic, uh, sources on the left, and then, uh, you know, Starburst in the middle and magic on the, on the, on the right side. Uh, so walk us through that. So in terms of, of sources, how do the data go into, uh, or how does Starburst access the data? What kind of data?
A Sure. Yeah. So at the heart of, uh, Starburst is this notion of connectors. And, uh, we think of everything as a connector. So even if you're just accessing S three and you're going to be querying iceberg tables, That's, that's technically our S three connector to our data lake connector, uh, to access that. Uh, and so every connector is basically just connecting to the underlying catalog of the system that you're connecting to. And then as soon as you've connected, which is like a, a one-time, you know, setup thing, uh, now you can run queries and you can do that at the command line, like just start to write SQL queries, joining tables across different systems, or, uh, you can use a BI tool like Tableau or ThoughtSpot or others. Um, or, and then this goes to the sort of data application side, we're seeing customers, you know, build more programmatic ways to interact with, uh, the data, um, that we have access to. Um, one of the features that we've built that's really nice, especially for, uh, internal purposes is, uh, something called data products, which is basically allowing you to stitch together a view of your data across these different data sources, and that's where you really start to get some, Interesting optionality because you can decide to materialize that view or not materialize that view, and, and there are trade-offs to both. You know, if you're querying the data…
AI assessment note: “at the heart of, uh, Starburst is this notion of connectors”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q But like, to which extent, in your opinion, does that help you take control over that?
A Yeah, I think this is a really interesting question. Um, I don't think it does allow them to take over the project, to be perfectly honest. Uh, I think that the market is actually resolutely determined to ensure that it continues to be independent, which is important actually for Iceberg, and that's what made Iceberg popular in the first place over Delta, you know, which was Databricks' you know, own format. The market wants an independent Standard. That's what they want, independent of any vendor. And so, fortunately, there's enough groundswell of people like ourselves, like Snowflake, uh, like some of the others you mentioned, where it is, and it's also, by the way, an Apache Software Foundation governed project. So, you have the Apache Software Foundation also ensuring independent governance, which is important and makes it truly, I think, independent. I mean, Ryan works for Databricks now, and he will for some period of time,
AI assessment note: “I don't think it does allow them to take over the project”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q So you started talking a little bit about the company. Do you want to finish that part? I guess, where are you as a company in terms of development? How many people? How long have you been around?
A Yeah, sure. So we've been around about 18 months. What's different about this one, well, actually a lot of things, it's funny. This company is both the same and very different at the same time compared to Hidapp. So Hidapp was You know, a sequel engine. This is again, a sequel engine. They're very similar from a technical perspective, but fundamentally the businesses are so different. This is obviously an open source project. We have very wide adoption. You saw some of the logos of some of the public users. It's a very global business. Um, the other thing that's very unusual about this one is, uh, we haven't raised any venture capital thus far. Um, and in fact, we're actually profitable, and we've been running, uh, profitably from day one, and And I will say we cheated, and I'll explain how we did that. Um, we were working at Teradata, as I mentioned, and we collectively left Teradata at the same time to start this business, and in doing so, worked out a deal that allowed us to continue to work with the number of customers we were working with around Presto. So we basically started with customers, started with revenue, and in that sense, this has been kind of like starting, uh, you know, with a head start, right, with a running start, and so that's been Really exciting. We're, uh, roughly 25 people today, uh, and growing organically thus far. Um, I won't say that we'll rule out…
AI assessment note: “we've been around about 18 months... We're, uh, roughly 25 people today”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Yeah. And when you say right, uh, what matters most is to like never lose any data and the data needs to be like a hundred percent correct. A hundred percent of the time. Exactly. Whereas analytical is, what do you optimize for in the, in the read?
A Yeah. So analytical is, um, uh, maybe a little less, um, I'll get, I'll call it maybe mission critical in the sense of, you know, losing a bite is not going to be the end of the world. That's the transactional side, which is very, um, focused on ensuring that, that consistency. But on the analytical side, you have a different challenge, which is how do I Process massive amounts of information. Do very complex joins of information across different tables, uh, to get to an answer as quickly as I can. And that requires a lot of, um, I'll call it rocket science in the optimization of how those queries actually get executed. And, and that's, that's what we focus on on the analytical side.
AI assessment note: “how do I Process massive amounts of information... to get to an answer as quickly as I can”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Uh, how do you think of, um, you know, a classic question, uh, for, for, For, um, you know, uh, open source businesses, like how do you think of that enterprise offering versus the underlying Trino, uh, functionality, open source functionality?
A Yeah. So I would call enterprise a classic open core model, which is to say that the core is open, meaning, you know, there's a Trino engine inside where the leading contributors to Trino, you get that as part of the offering. Uh, and then around it, we've built all the enterprise functionality that you would need. So fine grained access controls, you know, row level, column level, data masking, Query auditing. We've also built in a number of extra performance features. We have something called warp speed. Um, we have a lot of fun with the names of our sort of sub products. That is smart caching, smart indexing, uh, that delivers, you know, 10 X performance boost over faster SQL, faster SQL. Exactly. Exactly. Some, some techniques there that give you faster SQL. That's exactly right. Um, you know, we have extra connectors, we have, uh, you know, management capabilities, um, and so forth. And so that's, That's what Starburst Enterprise is. Because it is self-managed, what we mean by that is the customer is managing it, and that allows them the flexibility to deploy it anywhere. It could be deployed in an air gap facility. It could be deployed in a vehicle if it needed to be, like, you could run it literally anywhere. Um, and, and so again, that, that works for customers who have on-prem environments, hybrid environments, very complex environments, and gives them a lot of flexibi…
AI assessment note: “I would call enterprise a classic open core model, which is to say that the core is open”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q cheating because I'm, I'm bringing up the, the chart in front of me, but I'm not showing it to you. Um, so you have, joke aside, you have, you have, um, so ingestion we talked about, you have table maintenance. So governance we talked about, accelerated SQL analytics we talked about. The two things we didn't talk about is table maintenance and automatic capacity management. Do you want to go?
A Okay. Yeah. Yeah. So the table maintenance piece, that's, um, you know, there's a lot of activities that you want to do to Optimize for performance reasons. How the, the, the tables are, are structured and laid out. If you've been around for a while, you know, this is kind of like defragging a disk drive from like a long time ago, right? Like you want to take a lot of small files and bring them together into larger files, and that's called compaction. And that's a particular thing that you want to do with iceberg tables to get better performance. And you could do that yourself. It can be tedious, time consuming, error prone. There's a lot of work involved. Or you can use Galaxy and we do it all for you basically. So that's, that's sort of the table maintenance side of things. And then on the capacity management, that's really those auto scaling capabilities where you can spin clusters up and down and have them automatically scale based on incoming compute. It's a more serverless type of, uh, experience, um, behind the scenes.
AI assessment note: “So that's, that's sort of the table maintenance side of things. And then on the capacity management”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q So actually a very interesting, um, topic services for, you know, any founder listening to this. Like, how do you, how did you go about Building that network, so Kubrick or others. When did you feel it was the right time to start working with SIs?
A Ah, that's a good question. I would say I would say, you know, you should probably learn how to do it yourself first. So I would say not at the very, very beginning, but then my advice would be, um, try to start training up, uh, one or two and probably boutique firms. Like don't, don't go after Accenture on day one, uh, but start with a smaller boutique firm because what's most important, especially in the early parts of developing the services piece is the quality of the delivery. And you want to control the quality of that as much as possible. Obviously you control it when you do it yourself, but that's, that's difficult to scale. And it's, you know, it's, it's not the margins that you're, that your investors probably care about. They don't want you to build a services company. So that's probably not going to be the dominant, uh, way you want to deliver ultimately. But, um, by choosing just one or two and really focusing on them, there's a, there's a mutual benefit there. A, you're helping to bring them business and you're creating that positive reinforcement cycle that Learning how to deploy your software is going to lead to more revenue for them, but also you're, you're creating a reliable partner that you can count on, that, that you know, when you recommend that firm to deliver services at customer X, you know, they're not gonna make you look bad, you know, and that's, th…
AI assessment note: “not at the very, very beginning, but then my advice would be”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q So you are very early to that vision of the lake house, now ice house, uh, was it, uh, I don't know, controversial at any point, the, you know, uh, iceberg plus Trino as a reference architecture. Or was it hard to convince people that it was going to be the future?
A I would say yes. Until last year, there was a lot of debate over which format is going to win. You know, Databricks had created one called Delta. There was another open format called Hootie. Uh, and then there was Iceberg. And if you're just evaluating those for the first time, you think, well, there are three formats. How do I know which one's going to win? Um, I think the reason we had conviction that it was going to be Iceberg was simply that It was the one that had already been adopted by a lot of the super scaled up internet companies, and I think there's something to be said for saying, um, that, you know, you can watch those companies and kind of see where technology is probably going to go because they're the ones running at the most ridiculous scale. These, these technologies get really, uh, tested to the limit that way, uh, and can be a good indication of sort of where things are going. So we saw them all adopting Iceberg along with Trino, And, you know, felt like that was likely going to be a pattern that, um, is adopted by the industry, industry as a whole.
AI assessment note: “I would say yes. Until last year, there was a lot of debate”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Okay. All right. So that's Galaxy. And then we alluded to the Dell thing. That's more recent, right?
A Yes. That was almost 12 months ago that we released what's called the Dell Lakehouse. So it's, it's their product powered by Starburst. And that's essentially, you know, starburst inside with Dell's object storage, Dell's hardware, uh, and it's become a real centerpiece of their AI, uh, strategy in, um, selling this into, uh, large customers. Some of them themselves are CSPs that are, uh, building out their own, uh, infrastructure for, you know, analytics and AI. And I think at the end of the day, AI is only as good as the data that you're training the models on, only as good as the data that you're accessing. Through RAG workflows, and so having a lake house at the core of that architecture, um, you know, is, is important.
AI assessment note: “Yes. That was almost 12 months ago that we released what's called the Dell Lakehouse.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q That must have been a, was it an obvious decision or was it like a super nerve-wracking decision?
A It was, it was a nerve-wracking decision. It was obvious that, um, we had prototyped this, this different approach that became Galaxy as it is today. And, and it was obvious that that was going to be better, but even still, uh, it was a nerve-wracking approach because you're sort of throwing away like two years of development and starting over again. Uh, and you know, of course, we're venture backed, and we spent a lot of money to build that, and that was actually one of the reasons we raised venture in the first place. Um, you know, we have a little unusual history that we were bootstrapped the first two years, and we were running a nice little profitable business. It was, it was, it was great, but we thought, you know what, we can't build a SaaS solution. We can't build a cloud platform off our little bootstrapped, you know, small business. It's just, you know, too expensive, and so we raised venture, built this, Cloud platform and then, you know, threw it out and built it again. So that was definitely stressful.
AI assessment note: “It was, it was a nerve-wracking decision. It was obvious that”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Where do you fit in the, uh, sort of AI stack, uh, world?
A Yeah. So two places, I would say. First and foremost, on the training of models for those who are actually building their own models. Your models are only as good as the data that you train them on, and, you know, access to more data or better data is going to influence that. And so we become this sort of, um, access layer to all the data in our organization. There is no piece of data that we cannot get to. Essentially, um, through our federated architecture and being able to work across on-prem and cloud. And so that's one element is, you know, you're gonna access data, you're gonna do some transformation of data, prepare data, train a model with that data. The other is the more practical aspects of, uh, putting AI into production, which, you know, we see, uh, RAG workflows as an essential part of that. You know, people are building these agents that need to access contextual information Pass that along to the LLM to get the appropriate response as part of the, the, the, the agent or part of the application that they're building. And we think we can play a central role in that, uh, both in terms of, you know, the access to structured data, but also, uh, as, as performing, uh, vector search. And so we expect to be, uh, doing more, and you'll hear a lot more about us, I think, this year, uh, around RAG workflows and how Starburst plays a unique role in that.
AI assessment note: “two places, I would say. First and foremost, on the training of models”