Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 5 · Cm 5 5.00
Q And, um, maybe to, to, to put that in context further, just some of the, some of the names of, uh, you know, famous data warehouses, and a lot of folks are going to know this on, on, on, on, on the Zoom, but, uh, what are some examples of the main data warehouses?
A Yeah, the most popular data warehouses today would be, uh, Snowflake, who everyone's talking about right now because they're about to IPO, um, Redshift, Uh, which is a data warehouse that you can buy from Amazon Web Services. Redshift was incredibly important because, um, it was, it came out in 2013. It was in the AWS console. It wasn't the first really good fast data warehouse, but it was the first one that was cheap. Uh, and so a lot of people bought Redshift who previously would not have been able to buy one of the enterprise data warehouses that existed before that. And then Google BigQuery. Is another important data warehouse, uh, today that a lot of companies use.
AI assessment note: “The most popular data warehouses today would be, uh, Snowflake”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q And what happens after the data warehouse, um, in terms of analysis? I mean, they, there is, there's, I guess the BI world and there's a machine learning world. Like what, what, what do people do, um, sort of, uh, yeah, after the, the data warehouse in the, in the pipeline?
A So data warehouses are database management systems. And so fundamentally you can do anything you can do with data you can do with a data warehouse. Uh, in practice, the most common use of data warehouses is to support, uh, business intelligence dashboards. So these are dashboards. You've probably seen them if you've ever worked at a big company that tell you they have bar charts and line charts and things like that. They tell you what's going on, you know, How many support tickets were filed this week? How much, you know, how many bookings has the sales team done? What is the, you know, average response time of the website? Um, or if you're in, you know, the automobile industry, you might care about what is the average, you know, uh, value of our total inventory from our suppliers, which is something we're trying to minimize. It's always very business specific what your key metrics are. Um, but those are, uh, Generated from data in the data warehouse and then they're presented in a dashboard of a BI tool like Tableau or Looker or Microsoft Power BI or you name it. Um, so that, that's definitely the most common use case of data warehouses, but then you can really do anything with them. Um, we have customers who run billing out of their data warehouse. We have customers who, um, There's all kinds of use cases, uh, that can happen with data warehouses, because at the bottom of it,…
AI assessment note: “most common use of data warehouses is to support, uh, business intelligence dashboards.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q We talked about transformation, uh, a minute ago. Um, I see you, you have a product now, right? Your five trend transformations. What, um, is that correct? And, uh, I guess, how do you, how do you go about it?
A Yeah, that's right. So when Fivetran delivers the data, it's going to be in a normalized schema that is like a sensible schema. The data is clean. Uh, there's not like duplication of the same information across tables. Everything's up to date, but that schema is not going to be customized to your particular needs, right? It's going to be like the native schema of whatever the data source is. And for you to do anything useful, With that data, you're going to need to transform it typically into a dimensional schema is what most companies will do. They'll, they'll turn it into a dimensional schema, which if you've never heard of it, it's basically a, A simplified view of the data where you make some simplifying assumptions knowing what kinds of analysis you want to be able to support later. So everyone's going to have a different dimensional schema. Uh, and somehow you need to orchestrate this transformation, right? As a practical matter, an analyst is going to write a bunch of SQL queries that transform from one schema into another, but somehow you need to, like, store those SQL queries somewhere. You need to Keep track of them and review changes to them. Uh, and then you need to actually run them. Uh, and so in order to, uh, how to do this has been somewhat of a, like an open question for the last few years. We're not the only ones who have been pitching this ELT modern data sta…
AI assessment note: “Yeah, that's right. So when Fivetran delivers the data”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q the first round was in the 2018, right? And then you did multiple rounds, like back to back, like ABC sort of like in compressed timelines. So there was like a period of like five years, right? Where, where you were sort of, uh, two years building and then like, you know, several years into rating to get to that stage. Is that the kind of timeline we're talking about?
A Yeah, we, that's right. We raised a modest amount of money from angel investors right way back at the beginning. Um, so a few 100,000 dollars, and that's what sustained us for those first couple years. We did not pay ourselves a lot. Uh, and, uh, and we didn't spend a ton of money on AWS, and we only had one other person who joined the company in that, in that first phase. Um, The, uh, and then we raised, uh, the, the first significant round was a seed round in, in 2017, um, from a family office called CEAS, and then there was the series ABC in fairly short succession because we started to grow so fast. One of the funny things I learned about fundraising from that was that, um, if the company is growing really quickly, this funny thing happens, even if you don't spend the money, that same pile of money, Look smaller and smaller compared to the size of the company and the amount of money that just goes whooshing through every month. So no matter how capital efficient you are, you end up the, the faster you grow, the sooner you need to raise again, unless you want to just sit there and have, you know, one month's payroll in the bank account, which I don't think you want to when you have a lot of employees. So it is this funny paradox of fundraising, uh, that like the timing between rounds kind of doesn't matter. It's if it's a big Series A or a small Series A, you're going to do …
AI assessment note: “Yeah, we, that's right. We raised a modest amount of money from angel investors”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q You mentioned the word, uh, reverse engineering. Do you, presumably you have to partner with a lot of those, right? I mean, you have to, you have to work hand in hand with the source to figure it out, or is that something you can do with that?
A We, we do that more now. So in the beginning, I mean, we were on our own. We were small. Nobody wanted to talk to us. I mean, sometimes it was a battle just to get them to even give us access to the API. Uh, so in the beginning, we were just on our own. We would go and read the API docs. We would set up a test instance and do experiments, uh, play around, try to break it. Um, I would always joke that we were, we were just reversing The company's API just back to the normalized schema and the database underneath. Um, and then we would put customers on it and they would break things and they would call us up and be like, Hey, this row in my data warehouse doesn't match, uh, the, the source. And the really critical thing we did that laid the foundation for our eventual success was, uh, that we said that it was our responsibility to make the data match. Which is actually unusual in the field of data integration. Most data integration tools, they see themselves as like a platform, right? So they give you all this sort of toolkit, all these Legos and they say, yeah, but it's ultimately up to you to achieve correctness. Like you need to assemble this thing together and like, we're going to call the API endpoint and we're going to load the row. But if the data doesn't match, like that's your problem. Whereas we said the data will match like the schema will exist in your data warehouse …
AI assessment note: “We, we do that more now. So in the beginning, I mean, we were on our own.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q the, the new CMO is not a brand person. It's a data person. Uh, but are you finding that you need to do, um, A lot of, uh, sort of education, not about the need that they perceive, but about what needs to happen next. I mean, do those folks understand what a data warehouse is, what a connector is, and, and all those things? Uh, how does that work?
A Yeah. Usually by the time they're talking to us, uh, they know what a data warehouse is. They know what a BI tool is. Uh, and you know, Fivetran is arriving to solve Uh, a access to data problem that they need to solve in order to have their dream dashboard complete. Sometimes I joke that we're, we're sort of like the plumbers building the house. Like they go to the general contractor first, and that would be either like the data warehouse or the BI tool. And then like, they kind of work their way through the project and they're like, okay, well now the toilets need to work. Now we're going to hire a five train. Not that I'm diminishing the Uh, gloriousness of what we work on.
AI assessment note: “Usually by the time they're talking to us, uh, they know what a data warehouse is.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q And how does the, uh, automation park part work? Like, do you, do you ping the sources at regular intervals? How does that work?
A So it's different for every source. Uh, that's what makes it such a mammoth task, uh, to build connectivity to all these different data sources. For every single source we support, we have to go in and we have to understand how the API works. We have to understand how the underlying schema of that system works. And then we have to take, you know, we have to figure out a sync strategy for every source. We have to figure out a change data capture strategy where, well, you know, Call up this API endpoint and say, hey, tell me what changed, and then it'll give us a bunch of data, and then often there's weird rules about how, you know, this column represents the time it was modified, except when this happens, and then it's something else. Sometimes I joke that the real value of Fivetran is all the millions of if statements inside of our code that map out all of these scenarios, and so a big part of it is just elbow grease, working away on over all these different data sources, databases in particular, Are crazy. Um, I mean, you can make a whole company just out of syncing one database, and we sync, I think, uh, seven different major databases at this point. Uh, they have these very complex changelog formats, but then there is also a shared platform that we have internally. Um, so over the years, we've discovered a lot of fundamental principles of what makes connectors reliable, what…
AI assessment note: “we have to figure out a change data capture strategy for every source”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q And what's, uh, powered by a five trend that looks like the, um, Most recent product enhancement at least, at least I saw. What does that do?
A So powered by Fivetran, uh, is a way for you to embed Fivetran into your own application. So I've talked a lot about You know, data warehousing and business intelligence, uh, and, uh, there is, uh, another way to solve this problem, uh, which is vertical integration. So you can build your own data warehouse and you can hire a bunch of analysts who write SQL queries and then, uh, use a tool like Tableau or Looker to build dashboards that you present internally. Lots of companies do this. Lots of companies are going to continue to do this. However, it's a huge amount of work. Uh, it's incredibly expensive. Not so much because of the tools, but because of all the people. Um, and so you're gonna run out of steam. There's only so much you can do, uh, especially at a smaller or mid-sized company. And the other way to solve this problem is through vertical integration. So there have been for years companies that focus on data analysis in some domain. There's a lot of them in marketing, but there's also companies that focus on, you know, data analysis for, Consumer packaged goods. There's a Fivetran customer who focuses on data analysis for dentists, okay? Because you know what? There's like tens of thousands of dentists, and they're not going to hire analysts to write SQL queries for them, but there's still a lot of useful insights that they can derive out of their data. And so what p…
AI assessment note: “powered by Fivetran, uh, is a way for you to embed Fivetran into your own application.”
Answered raw tape
D 4 · C 5 · P 4 · Cm 4 4.30
Q And then presumably that needs to work in all environments, right? Whether the source is cloud base, on-prem, hybrid, um, Does that, does that work everywhere? Are there unique challenges to the different situations?
A So the API data sources we support are all, um, cloud-based apps. There are a few things like Jira you can in principle deploy on-prem. I don't know whether any of our customers are actually running Jira on-prem. It is, you know, less common than maybe it once was. Uh, and there's a few other data sources like that. Where in principle, they might be actually running on-prem, but there's no way for us to tell. It's just a URL. And then we support databases. We do actually sync a lot of on-prem databases. So we have a lot of customers where their production systems are still on-prem and the databases are running there, but then their data warehouse is in the cloud. Uh, the data warehouses we support are all in the cloud. They're always in the cloud. And the reason is the cloud-based data warehouses are just so much better. There is a reason why Snowflake is about to IPO for a bajillion dollars. And it's because- Cloud data warehouse. Yeah, exactly. I, I heard it was two bajillion. Uh, it's because cloud data warehouses are so much better, uh, at just even, I mean, they're a particularly good one, but then the category is just so much better than the previous generation of data warehouses, um, because if you think about it, you know, what is data warehousing all about? It's all about storing lots of data. Like I said a second ago, you want to store everything that ever happened in…
AI assessment note: “We do actually sync a lot of on-prem databases.”