Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Which was in, like, 2006, right? Yeah, 2004 to 20 years.
A 2006. Everybody's brain just sort of, like, broke, and they're like, wow, in order to build systems that can handle the data sizes that we're seeing, you kind of, you have to just Dramatically change how you're building them. You have to, you have to run on lots of machines, lots of cheap, inexpensive machines versus, you know, these giant hyper expensive machines. And to be fair, that was a problem at the time. Um, but you know, nowadays, like, you know, I've got a Mac M two laptop. It's two years old. Um, it's like probably an order of magnitude to two orders of magnitude faster than the server machines were back when, you know, MapReduce came out and people started building Building these systems, you know, let alone like nowadays the server systems, you know, have hundreds of cores and, you know, can have terabytes of RAM. Um, and, uh, and so really kind of like if you were going to build something now, like why would you bother with all the complexity of scale out? Because the thing about the way we designed systems, like I was one of the people that helped start Google BigQuery, uh, and I worked on, you know, single store for a couple of years. So I, um, you know, I, Have been in, you know, with my elbows deep in, in, you know, building these, these, these complicated systems is that, like, there's just this huge tax that you pay to, um, to have to build a distributed sys…
AI assessment note: “2006. Everybody's brain just sort of, like, broke”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q uh, if you use, uh, data as in, you know, modern data stack for purposes of BI, effectively data analysis, uh, then, uh, you know, maybe small data makes a lot of sense, but if you wanna, I don't know, train a big machine learning model, then, you know, as abundantly documented, you want as much data as possible. Is that, is that, uh, is there some truth to that?
A Yeah, absolutely. There, there are clearly some use cases and some workloads that, you know, that are not, are not small data. I think, you know, big, big data may be dead, but it's, you know, not, uh, it's not going away. Um, and, uh, I, I think the one, one thing that I often hear though is, um, you know, sometimes people are like, well, I read your, I read big data is dead and I, I, I, you know, I agreed with, with most of it, but, um, but I've got big data, you know, uh, but I, I think actually, Because there's a lot of people that have a lot of data, but what they actually use is a small section of data. So you might have 10 years worth of logs. It might be a petabyte worth of logs. Um, but if you only actually query the last seven days or the last day, um, that's, that's not really big data. Because of separation of storage and compute, that other data just sort of sits there and is culled on, you know, on AWS S three, Um, on, or on, you know, on object store somewhere, the only, the important part is the part that you're querying, and, you know, the vast, vast majority of workloads actually just query that, um, that hot data because it's, um, because it's expensive to, to query the whole thing. I mean, like, I used to run this query, um, I used to, I used to give talks on, on BigQuery, and I'd get up on stage, and I'd say, hey, look at this data set, it's a petabyte, and…
AI assessment note: “Yeah, absolutely. There, there are clearly some use cases and some workloads”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q So I'd love to go into, uh, Mother Duck specifically. So what is DuckDB?
A You know, DuckDB like, like, um, like SQLite is just a library. Uh, it's something that, you know, as you're building your code, you link that in. It's, you know, functions that you can call inside that library, and you don't have to set up a separate server somewhere and call out to that server. Uh, it just sort of runs everything inside, inside your, inside your process. And from a, you know, from a latency perspective, that's, that's super nice. From a complexity perspective, it's also super nice. You know, if you're running in Python, for example, um, You know, the, the way Python works is kind of things in Python have access to the, the variable namespace, and so actually in DuckDB you can query against your Python objects. So like, so if you just link in, you know, if you just import DuckDB actually in your, in your Python process, you can just start querying, querying data without even having to move it or prepare it or do anything. So it's just, it's just really, really kind of Very high developer, uh, experience, uh, for accessing your data. Uh, it also works, you know, well as a sort of standalone, as a standalone database, as well as not just sort of in memory, but it's, uh, um, you know, I think it's, and it's also, because it's just a library, it has no dependencies, and it's just, it's, it's incredibly lightweight. It can run in the browser, so it can run in under…
AI assessment note: “DuckDB like, like, um, like SQLite is just a library.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q guess the community does what the community does, but, um, you know, it's sort of, uh, What's the word? Like a sort of a common wisdom in startup and venture circles that the, you know, the, the commercial company building on top of the open source should also kind of own the community to the extent that any community can be owned. Like, how does that work for you guys?
A Right now we kind of have, you know, they're somewhat disjoint communities. We have a, we have our mother duck community and we have like a Slack. Um, but also there's a pretty vibrant DuckDB community and that we have not, You know, we have not jumped in and tried to own or run anything. I think, I don't think that would have gone over well. Um, on the other hand, we try to be helpful where we can, you know, we, we contribute a lot back to, um, you know, to DuckDB code. Uh, and, you know, we have great, you know, relationships with the, with the, with the founders and with, and with the community. And, um, and so I think that there's, I think that that's, you know, It's actually kind of a happy, uh, happy way of, of, of doing things. I mean, DuckDB is, is super popular, and, uh, you know, we don't want to try to, um, kind of horn in, horn in on that as long as, you know, we can be, you know, we can be successful by, you know, building our, building our managed service.
AI assessment note: “we have not jumped in and tried to own or run anything”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Does it impact you, uh, at all? Like, you know, in a world where, um, you know, tabular, uh, and, uh, you know, that was post the acquisition of, Tabular by Databricks and the rise precisely of those open data format. Does that change anything for you?
A It does sort of open up some doors for us. If you think about, um, you know, a lot of people who's, uh, we're hearing from a lot of people that, you know, that are using Snowflake and they're moving their data into usually iceberg, uh, as part of a, you know, as a cost reduction, you know, way to avoid lock-in. Uh, way to have, be more flexible and have more flexible, more flexible access to their data. And, um, and that's great news for us because if the data is locked in Snowflake, uh, and we want somebody to try, try Mother Duck, um, well, we have to convince them to, like, export the data or write it to two places, have multiple copies, and it's, it's a mess. It's a migration. Um, if the data is an iceberg, you know, we can just, we have the same access to it that Snowflake does. And, um, Um, and so I think as, uh, I think that's going to be hard for the incumbents, and I think it's going to be, it's going to be a net benefit to the, you know, kind of the people that are, people that are coming in with new, uh, with new tools and new ways of doing things, and it's going to put pressure on margins, um, which again is also, you know, tends to be in favor of, uh, of people that are coming in afterwards, especially if they have a simpler architecture and can deliver things, you know, less expensively.
AI assessment note: “It does sort of open up some doors for us.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q What about the commercial side of things, uh, coming from a very, uh, technical background? How did you learn there? What was surprising? Again, it was the, you know, a caveat mentioned up front, uh, that, uh, you seem to be doing an extraordinary job at, uh, marketing and, and marketing positioning.
A So I was an engineer for, for 20 years, and I kind of bounced back and forth with engineer and engineering manager, and I worked on a couple of projects that I thought were beautiful. I thought were like, wow, I'm just so proud of this, and that like, Are dead, you know, and they just, because, because like the commercial value wasn't there, the, you know, the company killed it or the company is no longer there. And, um, and so I think one thing that, that the lesson that, that taught me was like, you've got to understand customers. You got to understand the market. And so I was part of the team that helped, helped start Google BigQuery. And, uh, and at one point I ended up moving into, into product. Um, and part of the reason that I did that is because, uh, first of all, I was, Being asked to help hire the director of product for, for big query. And, uh, and I kept interviewing people. I'm like, wow, this, this person just really doesn't get what's special about big query. And like, this would be terrible if that person was, was, was, you know, led, led the, the product. Uh, and then I thought, well, maybe, you know, maybe I could, maybe I could do it, which felt like totally weird to, uh, to, to go from engineering to product. But that just opened my eyes in so many different ways, because as You know, the, you, in engineering, you're sort of, you're designing with this sort …
AI assessment note: “one thing that, that the lesson that, that taught me was like, you've got to understand customers.”
Answered raw tape
D 5 · C 4 · P 5 · Cm 4 4.55
Q and partners with, uh, George at, uh, Fivetran, uh, and I would have assumed that small data is not a good thing for the Fivetrans of the world, not to pick on them, but like any company in that stack, because don't you need a lot of data, uh, and a lot of, uh, complexity, uh, for, for, uh, this modern data stack, uh, this suite of vendors to thrive?
A I mean, it's, it's interesting you mentioned, you mentioned, uh, George and Fivetran because he, he, he was a speaker at our, uh, at our small data SF conference. He and I did a town hall conversation. And one of the things that he said was like, yeah, it's shocking how little data people, people use. And generally if people are pushing a lot of data through Fivetran, it's because they're doing something really inefficient where they're basically just sort of, they're doing basically, they're recopying all their data every day. Um, but, you know, typically, you know, they, the, the sizes of data that they see are, are much, much smaller than, um, than, than people would, would expect. But I do think that the modern data stack, you know, the ideas behind the modern data stack are important, and I think really that you have, um, uh, you know, you kind of have, like, I would call it sort of three, maybe four pieces. You have the, you know, data ingestion, Uh, you have kind of your query engine, and then you have your, your visualization layer. Uh, and then maybe the fourth would be kind of the orchestrator, um, and that would be sort of dbt. So, like, Fivetran would be ingestion. Snowflake would be the query engine. Um, Looker, the visualization layer, the visualization layer, and then dbt, the orchestration. Um, you know, obviously you can swap out those, those pieces, and there'…
AI assessment note: “it's shocking how little data people, people use.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 3 4.45
Q And, uh, so who created it? Who are the creators of it?
A So yeah, it was, uh, Hannes Mühlheisen and, uh, Mark Ron Ross, Mark Rothfeld. Um, uh, Mark, or Hannes was a, um, a professor. Mark was one of his graduate students, and Hannes had just gotten tenure, and so, like, kind of nobody could tell him what to do for a little while, and, uh, Mark had finished his PhD papers early, And so, nobody could really tell him what to do for a while, and they're like, hey, we've been using MONADB, and kind of, there's a bunch of limitations, uh, let's, let's write our own, and to solve some, like, some problems that they'd seen in kind of the data science world, and that, you know, they, they're like, data scientists hate databases, and partly it's because they had to install databases and configure them and load data into them, and they, they said, hey, you know, there's a, there's a, there's a, there's a better way, and it turns out that they're amazing, You know, they're amazing database researchers, but also, you know, great engineers, and they were able to build a, a super useful system that just started getting, becoming more and more, more, more and more popular.
AI assessment note: “it was, uh, Hannes Mühlheisen and, uh, Mark Ron Ross”
Answered raw tape
D 5 · C 4 · P 4 · Cm 4 4.30
Q is real time processing, which has a different stack. And then there was big data and the small data and sort of what do I do? It feels like things are getting more complex rather than less. Uh, but it's part of the small data message, uh, uh, that, uh, things actually getting simpler, but you sort of need to get rid of some of the Pieces of the past.
A I think one of the things, one of the things with small data is, uh, I think because the, um, because the architecture is simpler, we can focus more on building better experiences. So, um, yeah, maybe there, there might be, you know, a plurality of tools involved or a plethora of tools involved. Um, Um, the, if those tools are simpler to use, then, um, and I think kind of the net cognitive load can, can go down. And I'll give an, an example is, um, uh, a lot of data is in CSV files. And as much as sort of like as a database person, it sort of makes your head explode. You're like, why would you put all this in a CSV file? Um, it, it's just, it, That's the way the world works, and like, it's simple, and it's easy to, it's easy to write a CSV parser, and it's easy to read, to write CSV, but there's so much broken CSV out there, because it's actually really hard to write a totally unambiguous, you know, correct CSV file, and everybody sort of does it differently, different like null characters, and is, you know, two empty quotes, is that a null, or is that just an empty string? Like, there's just all sorts of, um, All sorts of weird, weird, weird things that happen, and, um, one of the things that DuckDB did is they said, okay, we're gonna, we're gonna really, really solve this problem, and we're gonna make it so you can just do select star from CSV file name, and it will do the ri…
AI assessment note: “because the architecture is simpler, we can focus more on building better experiences”
Answered raw tape
D 5 · C 4 · P 4 · Cm 3 4.15
Q And, uh, tell us about the company today. So I think you've, you raised about a hundred million, um, Across three runs, is that the right number?
A Yeah, we raised, you know, we raised in, um, you know, we raised our seed round from, you know, Redpoint and, um, Madrona and Amplify, uh, and then we got preempted a few months later, um, Andreessen, uh, and then, uh, we got preempted for RB about six months later for, uh, by Felicis. Um, and, you know, building a, you know, building a database as a service is expensive, You know, DuckDB is an amazing, amazing, amazing piece of software, but it's not a data warehouse. And kind of to turn that into a data warehouse, you know, is, is hard. There's a lot of things that kind of like you poke it, you poke it the wrong direction, it'll fall over. Uh, and then it's sort of like, well, no one's ever poked it in that direction before. Um, and so, you know, it's just stuff to stuff to work, work through stuff that, you know, having, you know, have to build and we're building this, I think pretty, pretty rich, Um, database as a service and the serverless, serverless backend that's highly multi-tenant, um, this hybrid execution system or actually dual execution system where we can push workloads down to the end user and kind of split query plans, and so I think we're doing some, you know, we're doing some interesting non-trivial, non-trivial stuff, partly because, you know, you've seen, there's a model that I, I think I've seen before, which is, uh, in open source where, you know, somebod…
AI assessment note: “we raised our seed round from, you know, Redpoint and, um, Madrona”
Answered raw tape
D 4 · C 4 · P 4 · Cm 3 3.85
Q What makes, uh, DougDB and MotherDuck then so fast and so appropriate for those use cases?
A You know, one of the things that makes it fast is it's just, it's, um, it's a brand new database built from, built from scratch, kind of applying, you know, kind of the latest and greatest best practices. Um, there was a paper that Michael Stonebraker wrote, um, uh, I think about 50, Michael Stonebraker's a Turing Award winner, like the Nobel Prize of Computer Science, uh, in databases and, uh, and creator of many database companies. Yeah, he created Postgres and Ingress, um, And I think the data, I think the paper was called No Free Lunch, and it was really about that, like, hey, the technology has changed, um, dramatically. Why haven't databases changed? Like the way, if you just think about, you know, um, SSDs, you know, like you might have in your, in your laptop, they, they don't have the spinning platter anymore. They're, um, um, and they're, because of that, they're much, much faster to find things. There's some things that they do that are better, that are dramatically better. There's some things they do that aren't, aren't as good. Um, but if you were going to build a new system, you would just build it differently. And, um, but people ended up like, well, the old one works kind of well, and like, and, you know, a decade, you know, it's a decade or even more since that paper, more has changed in hardware, in, in sizes and speeds, like you get these, you know, um, what …
AI assessment note: “one of the things that makes it fast is it's just, it's, um, it's a brand new database built from, built from scratch”
Answered raw tape
D 3 · C 4 · P 4 · Cm 3 3.55
Q And do you want to explain what in-memory means?
A So transactional database, you know, you're, you're typically, you know, you're operating on sort of one thing at a time. You have an order, and you, you, you create an order, and maybe you update the state of that order, and there's a bunch of, like, consistency checks to make sure that that order You know, matches a real customer and matches a real product and line items, et cetera. Um, and, um, analytical databases, on the other hand, tend to operate across, across data. So you can ask questions like, well, how many orders did I have in the last, in the last week? Or how many orders, you know, broke, who is my, uh, you know, biggest customer by, you know, by amount of data, amount that they spend, uh, and then broken out by region. And, you know, those, those, those types of questions, um, The data tends to be stored differently in these types of databases, um, you know, column store versus, versus row store, and, um, you know, and then there's also a subclass of databases, which is sort of an in-memory database, which means that if you turn your, if you turn your computer off, you know, you learn, you lose all the data, um, but on the other hand, memory is, you know, four orders of magnitude faster than, than disks, uh, depending on what kind of, kind of disks, but Um, it's, you know, generally much, much faster, and so you can do things, you know, s, you know, blindingly f…
AI assessment note: “in-memory database, which means that if you turn your computer off, you lose all the data”