Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q And I think with this initial positioning, you talked about the power of constraint on the business and how that creates the discipline that you need. Do you want to elaborate?
A Yeah, absolutely. And I think this is very universal. Whenever you're trying to, uh, attack a hard problem, sometimes it's great to really List out the constraints that you want to apply to that problem. So, you know, in our case, we had an up and running business, and we wanted to build a better company. So what is, what does that actually mean? And, ah, we said we wanted to have a monthly recurring revenue subscription model, because that would build a great revenue base for the business. We wanted to have a unique position in the market, and that was extremely difficult to figure out Uh, how do you enter the cloud market that was essentially dominated by AWS? And it turns out simplicity was that, um, kind of differentiation. We wanted to stay within the domain expertise that we had around networks and servers and just the data center. And, um, Feel like I'm forgetting one other thing, but the point is, it's counterintuitive in terms of you're really placing these constraints, and so solving for a smaller problem can actually, ah, lead to a more abstract and kind of general solution that will scale much better over time, rather than making something very complex and specific to a, you know, a, a point in time, and then what typically happens is, as that solution starts to scale, it can actually break down And won't really deliver on the test of time. And, you know, it's reall…
AI assessment note: “using these constraints, we were able to build a business that's extremely scalable.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q And the, the positioning was around the lower tier of the market in terms of size. Is that correct?
A Well, we definitely wanted to be an entry level place for, uh, creation for these, uh, businesses and developers that are building something new. And so obviously naturally you're going to start from, from nothing. But the good news is that we've been able to scale with our customers. Some of our largest accounts are spending over A million dollars a year with us. They're using thousands and at times tens of thousands of virtual machines plus, you know, a slew of other services that we offer today. But the idea was to help that initial moment of creation get people back on their path of building their application, building features and functionality, rather than trying to figure out how to manage their servers and their infrastructure.
AI assessment note: “we definitely wanted to be an entry level place for, uh, creation”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Yeah, and at a very practical level, do you want to talk about some of the hacks you use for bootstrapping, you know, server leasing, deferred payment plans, that, that type of thing, debt, any, any useful thing that people could use?
A I mean, just figure out how to use other people's money, really. A great, a great example. Uh, we had our launch, um, kind of event at the New York Tech Meetup here in New York, another big event similar to this one, and there was a law firm there, Gunderson Depmer, that was a sponsor of, of the evening, and they basically said, hey, come pitch us your startup idea. If we like it, we'll provide you with all the legal services, and we'll defer payment until, you know, in the, in the, in the future. And so, I had no idea, by the way, how much legal work you actually need to do, but I took them up on their offer. They thought we had a good enough idea, and so we started working with them, and I'm super thankful we went down that route. So, I think the idea here is, you know, how can you find sources of capital that aren't your own, and so, you know, debt is another really good, um, example where, You know, we, we've raised more debt than, than venture capital and venture. We've raised one hundred and twenty three million. And I think we've raised roughly like three hundred million in debt because we use, and we still to this day use debt to buy all of our servers and storage, um, and, and networking equipment, all the capex for the business because it just, it's that much cheaper rather than taking on the dilution and bringing on an equity investor. So I'd encourage everyone to re…
AI assessment note: “we've raised roughly like three hundred million in debt because we use... debt to buy all of our servers”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Maybe talk about the, the history of it, because it's been around for a while, right? When, when was it created?
A I don't remember off the top of my head, but it's a logical evolution from Bigtable, which was the previous storage system of choice at Google. And where it came about is Bigtable is only run in a single data center, and if you need to have more than one copy of your data across different data centers, which is probably a good thing, um, that's where Spanner comes in. Spanner is able to replicate your data across multiple data centers. The other switch that happened in Spanner is that we moved to a relational model, and what we found is that it's much easier for our application developers who are used From industry to using a SQL model to then use a SQL model on Google as well.
AI assessment note: “I don't remember off the top of my head, but it's a logical evolution”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q And, and the firm you came from before Third Point, World Quant is, is much more on the sort of pure quant side, right? Is that correct?
A Sure. So World Quant was fully systematic, which was more of a thesis of, if we consume more data than anybody else in the world, we can find more signals, create more alpha. I think on, you know, where we sit now, or where I sit now with Third Point, our team is thinking, there's all this data out there, and there's so much noise, we have to decide where do we spend our time, and what can help us better understand the names we have, and Rather than mining data for new ideas, we're almost looking for data to help us understand the ideas we may already have, or to, you know, create very sophisticated screens, you know, looking across all these companies to kind of dwindle down a smaller list of names that we can work on with a fundamental manager.
AI assessment note: “Sure. So World Quant was fully systematic, which was more of a thesis”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Do you want to explain what that is, layer one and layer two?
A Yeah, so, um, essentially, Um, web technologies have enabled us to do hundreds of thousands or millions of transactions per second for applications, um, because we have this new trust architecture for IT systems. We've kind of thrown away all that scalability and replaced it with this system that can only do 20 transactions per second, um, but you can trust, um, that that system can't be cheated, um, because it's radically decentralized. Uh, so now we need to At this layer, uh, at the layer one, we need to scale that significantly. Um, and that's going to happen in sort of phase three, but phase two, which is upon us right now, uh, is the use of off chain technologies, things like state channels where you coordinate and then finalize effectively an infinite number of transactions, um, or side chains, uh, where you can stand up, uh, a blockchain system fit for purpose, whether it's for a game or a Or a decentralized exchange, and you can link that into layer one, so that any digitally valuable elements that are in the system can be pulled into Ethereum, into layer one, and so maybe it's a digital sword that you pull out of one game, and you put into another, or you put that sword on an exchange, or something like that. So, there are lots of games that are on Ethereum right now, or that are coming out on Ethereum, that are essentially bringing their own scalability technologies, …
AI assessment note: “layer two, which is upon us right now, uh, is the use of off chain technologies”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q And just as a quick detour, do you want to talk about that, um, that, that, that post that, like, great innovation will look like a toy? I think that for anybody that's an entrepreneur, I think it's a super interesting concept.
A Sure, yeah, it's, um, it, it, I mean, it's really, it's not, I mean, I guess that was my blog post. The idea really comes from Clay Christensen Um, and, uh, the idea is, uh, is essentially that, um, like, if you go back, that essentially, um, a lot of technologies start off kind of half-baked, but they get better at a, at a rate that's sort of faster than kind of people need the technology to work. So human demand is sort of a straight line, and the technology is kind of coming up on a curve. And so, like, I remember Skype as an example, um, you know, when it started off, I actually worked at a VC firm. I was a junior person at a VC firm that invested in it. And in our memo, like, the key risk was there weren't microphones and computers at the time, really, literally, that was, like, the key risk, because that, you know, it wasn't clear they were going to be. Also, it dropped calls a lot. The quality wasn't that great. You know, if you looked at the time, all the business world was talking about VoIP and like these other kinds of things, and these sort of high-end experiences, but what obviously happened is it got better and better, and computers got mics, and bandwidth got better, and all these other things, and then, you know, and then eventually you had smartphones, and of course now, you know, that's sort of the way things work, so, um, there's just sort of, you know, tons …
AI assessment note: “a lot of technologies start off kind of half-baked, but they get better”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Great. Alright, so let's jump into use cases. Why do customers use you and how? Is it a consolidation type scenario of several data marts? Um, you know, is that a data lake type scenario?
A Every story with the customer is different. I mean, I, I think you see customers coming from different places. You know, we have some customers that Are, are trying to make Hadoop work, and are struggling with that to analyze machine generated data, and so they come to Snowflake from the machine generated side, and they use us for, for analytics associated with that, and, and those customers tend to think of us as kind of a big data solution. Um, they're, they're pretty much in the minority though, I'd say there's quite a few of them, but, but most of our customers come to us from some sort of, of relational data warehouse that they have, Which they, they want to have, they, which they, they are having some set of challenges with. Typically associated with the business teams not, not getting the performance or concurrency that they need, or the fact that the data is not able to be consolidated within a single system because of limitations associated with it. So they come to us from area, from, from those two different directions overall. Um, at first we saw more people who were already experienced with the cloud, And we're looking for the scalability that Snowflake can offer and the concurrency. So our first initial customers were really customers who already had experience in the cloud. They might already be running a, a database, a, a data warehouse in the cloud, and they wer…
AI assessment note: “most of our customers come to us from some sort of, of relational data warehouse”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q So the Hadoop and Spark ecosystems are friend or foe? Are they repositories that feed into Snowflake?
A I think they're different. I think they're very different. I think what we're seeing is, is a movement of the open source community towards technologies like Spark, um, which are, which are not storage-oriented, um, in, in their history. Uh, Hadoop, Hadoop's history with HDFS very much makes it a storage-based system, And then the, the, the, the typical approaches that people have traditionally used with Hadoop, variations of MapReduce, um, are now being seen, I think, as, as, now that there are alternatives, such as Snowflake that are available, that allow you to work with large amounts of data, and to do so with a true relational database, I think many customers who have previously tried Hadoop are moving towards a solution like Snowflake. So I, I think Hadoop is, is, is a, is a past technology. I think it's, it's, it's, although it's still gonna, people will still use it, it still has a place, I think it's, it's not an area where there's gonna be a lot of, of incremental additional investment. Spark is different. Spark, I think, is being used very, very broadly for advanced analytics, machine learning, in some cases for streaming data, and those scenarios are all very, very complimentary to Snowflake. We have a lot of customers We have a Spark connector, and we have a lot of customers that are running Spark in combination with Snowflake, and we think that's a pretty common s…
AI assessment note: “those scenarios are all very, very complimentary to Snowflake.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Great framework. Thank you. And a little bit to the last point you made. You hear some investors talk about, um, horizontal AI versus vertical AI. So horizontal being enabling AI that targets broad use cases, vertical being industry specific. Does that matter to you, or would you invest in a horizontal AI company?
A Yeah, it matters a lot. Um, you, you said this really early on, and from the very beginning of our fund, we just said we're not going to invest in anything that's horizontal. Um, And the reason is fundamental, right? Like, if you understand how to build a machine learning model, all the fun is in, like, tuning it for its specific purpose. Um, and not all the fun, all the value is in doing that at an algorithmic level, um, but also at the data gathering level. And we, so that's the first thing we thought, um, To, to make us only invest in vertically focused applications, because that's where you can really get ahead of everyone else by focusing on tuning a model for a very specific purpose, getting data to train a model for a very specific purpose. Um, so that was on the one hand, like a fundamental understanding of how this stuff works made us think you have to be vertical. The, on the other hand, uh, it's very clear that with this huge shift to cloud, as we just saw, like we're still only halfway through this shift, All these massive companies that are the cloud utilities, I call them, ah, or cloud infrastructure providers, they want to get all these machine learning work, machine learning workloads right onto the cloud because they're really data and compute intensive. So they have a huge incentive to give out whatever tools they can for free to get these workloads into their…
AI assessment note: “we just said we're not going to invest in anything that's horizontal.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Do you want to give a few examples from your investments of companies that have a unique or interesting data asset?
A Yeah, for sure. I'll start with a, something that's pretty simple and you would have thought is already done, but, uh, but hasn't been. It's a company called Constructor, and they give you a search, search box. That you can, you can put on your website, right? If you've got a media site, e-commerce site, whatever else. Now, Algolia, Elastic, like, all these companies do that. They make it very easy to deploy a very fast search box on your website, but the thing is, because they guarantee you that they're not going to share any data in any way with anyone, um, they can't really improve that search function over time in terms of, like, the autosuggest results, um, the ranking, Um, of the results. Once, once they're sort of shown up, showed on a page. Constructor pulls data across all of its customers. So it's got a bunch of e-commerce customers, media sites, and massive ones as well, like Jet.com is, is a customer of theirs. And it pulls all this data about what people are searching for, what typos they make, um, what time of day. If they're searching on mobile, do they have shorter, longer searches? What does that mean? What are they trying to find? Use all this data to figure out, like, what people are trying to do. And then look at all the click-through rates on all these sites, and, and provide a really, really accurate self-learning search engine. Um, you know, a lot of the,…
AI assessment note: “Constructor pulls data across all of its customers.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q So, um, you've been investing in those companies for a little while now. Any lessons learned in terms of what works, but actually what doesn't work?
A Yeah, um, just going off that example actually is, it, it's a good follow-on question, because our main lesson is, you've got to get customers to buy into the data network on day one, and if they don't, you should reject them as a customer, and so what I mean by that is, you have your terms, and your terms are, if you use our product, we will aggregate the data, and we will build models on top of that data, and we'll share the results or the improvements That improved model will be used across all of our customers. We're not going to share your specific customer data with another competitor of yours, but your data is going into a pool, and that pool makes the model better. So they're in your terms from day one. And your customers should either sign up to that or not. And because a lot of customers will say, well, that's scary to us. We don't want anyone seeing our data. And that's not exactly what's happening, of course. It's aggregated, anonymized, whatever else. But they might be scared by that. And some of those customers will ask you to do an on-prem private deployment, will ask you to not use their data to train anything that you do, and you need to reject those customers. Um, because you are not going to build a company that has any sort of moat around it, that has the world's best model, unless you get their data to do that. Um, and this is sort of like, ah, you know, 10…
AI assessment note: “our main lesson is, you've got to get customers to buy into the data network”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q So let, let, let's dive into case text a little bit. What, what do they do?
A So what case text Does, um, let me just give you an example of what it can do. So let's say you're a lawyer and you get a brief from opposing counsel and your client's getting sued. You basically upload that brief into case text and it will tell you every single case that, you know, it'll cite the cases and all the cases that opposing counsel missed, which more than likely are the ones you're going to want to cite, right, in your case. And they also, it was also crowd-based and crowd-generated, so they took the cases and all the citations and the references into those Cases were done really by crowds of students to do what we call a shepardizing in the, in the legal world, to make sure that the reference that they're citing to is the most valid thing that you can actually cite to. So, you know, if I was an attorney today, there's no way I would go to court or, or file a brief without checking it through case text, because as an associate, it makes you look like a rock star to your partner. You know you didn't miss anything, and as a partner, you know your associate didn't miss anything, and you won't Uh, you won't, you won't, you know, sort of be, you know, be disbarred, which in, in law you, that really is a concern. Like, you miss some important things, and unlike other professions, you can actually, like, lose, lose your license. So, um, so CaseX is really great, a phenomena…
AI assessment note: “You basically upload that brief into case text and it will tell you every single case”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q No feedback. All right. Uh, thank you for telling me that. So, uh, you, you alluded to some of this, but could you, um, maybe tell us more about AltaScale, what you guys built, and the story of the company, and all the things?
A Yeah, well, so my, my role at AltaVista and later on at, at Yahoo was to build more of the systems infrastructure than to do the machine learning. I knew enough to be dangerous as far as, say, relevance science was concerned, but, but my job was to take, The sophisticated algorithms and implement them at, at, at scale, and then to build the, the data infrastructure that the data scientists used to, to actually run these experiments that we talked about. So I did that, you know, again at AltaVista later on at Yahoo. Um, and, you know, Hadoop at Yahoo became kind of our standard infrastructure, and we made a large, large investment in that. Um, you know, through the spin out of, of Hortonworks, you know, that I was involved in as, as the CTO at the time, I kind of, We got a glimpse at how larger non-internet companies were struggling with the adoption of Hadoop. Um, you know, their clusters were kind of subscale. The, the people operating them were kind of operating them, and 10 other things at the same time. They'd have two people, not a hundred, using it, so if they, anybody got stuck, they'd immediately have to go search the web, you know, Stack Overflow to figure out what these stack traces mean. Um, and that, that was hugely unproductive. So, you know, at AltScale we said, hey, let's, you know, let's deliver Hadoop as a service And the way it's experienced at a Facebook or Y…
AI assessment note: “at AltScale we said, hey, let's, you know, let's deliver Hadoop as a service”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q The demo that actually worked in the, especially through the, World-renowned, uh, Bloomberg firewall. Very, uh, very impressive, among other things. Um, what else the, I mean, you mentioned a broad surface area. What else does a product do that you can talk about?
A Yeah, so the way we think about the, um, I like to describe the analytical life cycle, which I view as going from sort of early exploration and ideation, and so we support a number of interactive workspaces, like you saw briefly, Jupyter Notebooks, RStudio, Zeppelin. So spinning those up on remote More powerful hardware through our reproducibility engine that I showed you. The next phase I, I think of is sort of experimentation, and so that's mainly what I showed here. How can you run a lot of experiments, keep them tracked? And then the final phase is productionization or operationalization. How do you take what you built and get it exposed out of the business? So we support that in a few ways. You can deploy models as APIs for integration into production automated systems. You can wrap models you built in lightweight web forms, um, so that human consumers can interact with them without bothering a, a quant. Uh, we support kind of app hosting, shiny, Flask apps, things like that. And, um, and so that's sort of the life cycle of a particular piece of research. Then, you know, sitting underneath that is all this stuff around collaboration, preserving organizational knowledge, and so, um, I think I briefly touched on that. Commenting, discussion, search, knowledge management.
AI assessment note: “we support a number of interactive workspaces, like you saw briefly, Jupyter Notebooks”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Very cool. So any, any, uh, stats on Slack and how it's taking the world, starting with this room, apparently?
A Yeah, I mean, I guess this kind of sample is a kind of nice thing to see. Uh, you know, Slack, yeah, it's a very fast-growing company. I think, ah, our public stats are that we have passed over three million daily active users, so people are using the product literally every day. Ah, those daily active users are using it for about two hours, over two hours a day, like, actually foregrounded on their phone or on their desktop. Ah, they send over 60 messages a day. We have billions of messages sent every day. Ah, so, I mean, quite frankly, that's one of the reasons why we all, ah, we're excited to kind of join to work at Slack at this point, where Uh, there's definitely enough product market fit and enough kind of intensity of usage to start really doing interesting things about understanding, uh, the work graph and how people relate to each other and how people relate to channels, uh, and also kind of understanding, uh, all this unstructured information that's passing through Slack just so people can get their job done. Can we start making sense of that? Can we start organizing it, ranking it, and feeding it back into people, uh, to make their, you know, working lives more productive? Uh, so that's, that's kind of the moment in time, uh, that, Kind of bore out this group that we're starting here in New York, which is, yeah, like you said, called the Search Learning Intelligence …
AI assessment note: “our public stats are that we have passed over three million daily active users”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Very cool. So any, any, uh, stats on Slack and how it's taking the world, starting with this room, apparently?
A Yeah, I mean, I guess this kind of sample is a kind of nice thing to see. Uh, you know, Slack, yeah, it's a very fast-growing company. I think, ah, our public stats are that we have passed over three million daily active users, so people are using the product literally every day. Ah, those daily active users are using it for about two hours, over two hours a day, like, actually foregrounded on their phone or on their desktop. Ah, they send over 60 messages a day. We have billions of messages sent every day. Ah, so, I mean, quite frankly, that's one of the reasons why we all, ah, we're excited to kind of join to work at Slack at this point, where Uh, there's definitely enough product market fit and enough kind of intensity of usage to start really doing interesting things about understanding, uh, the work graph and how people relate to each other and how people relate to channels, uh, and also kind of understanding, uh, all this unstructured information that's passing through Slack just so people can get their job done. Can we start making sense of that? Can we start organizing it, ranking it, and feeding it back into people, uh, to make their, you know, working lives more productive? Uh, so that's, that's kind of the moment in time, uh, that, Kind of bore out this group that we're starting here in New York, which is, yeah, like you said, called the Search Learning Intelligence …
AI assessment note: “our public stats are that we have passed over three million daily active users”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q You know, any one or two that come to mind in terms of doing interesting things?
A Yeah, I mean, there's a bunch that we see kind of in the Slack ecosystem that we're kind of excited about, um, from things that are more technically simple, but I think are kind of delightful, like, uh, there's a company called Growbot that is building kind of an app to give kind of positive reinforcement within the workplace. Uh, there's a company in New York that we've invested in through our Slack fund, uh, just raised another round a couple weeks ago called Troops, which is kind of building Like an intelligent CRM that's a kind of conversational interface to start with. So there's a lot of interesting things. The other thing that we see a lot is at larger companies, we certainly do this at Slack, people write kind of their own custom integrations to plug into their back-end systems, so we do this for pretty much every stats system that we have is a bot that kind of emits regularly different kinds of alerts, and I think the companies that wind up loving Slack the most, they basically take all these really ugly internal tools that they have that no one ever wants to log into anyway, And they start just kind of connecting those pipes into Slack, uh, and that's something that winds up being kind of magical experience for them.
AI assessment note: “there's a company called Growbot that is building kind of an app”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q earlier in the conversation that, um, Slack at the very beginning was a distributed company before it became distributed, and so Slack is headquartered in San Francisco, but you're based here, and you're building this team here, so is that, is that circumstantial? Is that like, no, he's this great guy, let's have him, and it'll be the team in New York, or is there something special about New York?
A Uh, it's a good question. I would say there's probably historical, corporate, and personal reasons for why we started up the team in New York. Uh, I mean, one historical is, yeah, Slack started off, there was four co-founders. Uh, one of the, the first person to be in New York actually before I joined was actually Eric Costello, who is effectively the tech lead for the front end of Slack. Uh, he's worked from New York for the 15 years since before Flickr. Uh, and two other co-founders including the CEO were in Vancouver, and one who's now the CTO, Cal Henderson, is in San Francisco. So, one, it's kind of path dependent a little bit. Slack only exists because they have this distributed team. Uh, I would say the corporate side is, you know, many of the biggest customers that, uh, already exist for Slack or will exist in the future all exist in this New York City kind of region or the Northeast region. Um, quite frankly, based on the survey in the room, you can probably tell we've reached almost saturation in San Francisco and the Bay Area in terms of companies using Slack. Uh, so most of the, I think the growth will happen in other parts of the world, and this is kind of the economic engine of the rest of the country. Uh, so we were always going to kind of come up with an enterprise office here. Uh, but I think what we saw, what I saw having worked at Google, worked at Foursquare…
AI assessment note: “historical, corporate, and personal reasons for why we started up the team in New York”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q is the Internet of Things, um, and it seems that insurance companies are the, very closely related in terms of, um, actually doing something with all the data exhaust that's coming from all those, uh, connected devices. So, maybe, maybe any additional details on telematics? Is that, is that, is that some specific hardware that you guys install in cars? Is that a mobile app? Do you work with the
A Exactly. Well, it's a partnership. And so, uh, currently we're using, uh, specific devices that we'll go ahead and attach into a car's, um, uh, maintenance port. And so we can go ahead and we can milk information from the car, speed, geolocation, all that information. And that, again, it helps us to look at patterns as to where people are driving and when they're driving in those areas. Is it a congested area? Are there a lot of stop signs or, or, um, are there a lot of children in those areas? And so all that information, It helps us to really understand from a pricing standpoint, ah, and the risk standpoint to really make sure that our customers are having a, a good experience and, and getting the best, um, benefit from, from our products.
AI assessment note: “currently we're using, uh, specific devices that we'll go ahead and attach into a car's”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Great. Um, you alluded to this a little bit, but tell us a little bit about the, I guess the data, The tech stack, the data infrastructure, what, what do you guys use? Is it, and particularly in the context of a gigantic company like AXA, does, does US have a different tech stack than the rest of the world? How does that work?
A Well, because of a lot of the regulatory environment, one of our challenges is that data can't flow freely from continent to continent, ok? Ah, and so one of the things that we've done is we've essentially built three main, ah, data Analysis hubs around the world. So we've got one in, uh, the, uh, France, uh, in, in the Paris location. We've got one here in, uh, in U.S. to support all the Americas, uh, and then we're, uh, in the process right now of building one in Singapore as well. Um, so the one that we've got here in the U.S., it's, uh, it's primarily a Cloudera Hadoop stack, uh, that we've, uh, uh, used a blueprint that was essentially blessed by our, our brethren over in French, in France, Um, and with that, we've got R and Python and Spark. Uh, we use some Dataiku along the way, and so all these products kind of come together to give us a, a capability for not only our actuaries, but also our data scientists to, to access the data that they need.
AI assessment note: “it's primarily a Cloudera Hadoop stack... R and Python and Spark”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Yeah, no, that makes perfect sense. Um, so tell us about, uh, your investment is just, I guess, what, what gets you excited these days?
A Um, I, I guess I can, I can put a, um, I, I can put a data angle. I was, I was thinking about this from a data perspective. I mean, there's a lot of things we're investing in at Trinity. Um, there, there's really, um, you know, I would say two themes that we're interested in when you think about, uh, companies that are doing things with data. I think one of these has, um, been, already been stated, but it's solve, solve problems, right? Um, with whatever, with whatever data you have. Um, at, at Logly, which is one of my portfolio companies that does log management and analytics, uh, as a cloud-based service, so think Splunk in the cloud, we talk about revealing what matters. And so, you know, what, one of the things that we're looking for in, in data companies is not just handing the user a bunch of data and saying, hey, we've got tools for you to figure out what's interesting here, but actually companies that reveal the in, That, that, that pull the insights out of your, uh, out of your data and make that very easy for business users to understand what's going on, um, without having to understand complicated query languages and things like that. So Logly is an example in, uh, uh, the log management space. Instead of logging in and just seeing a blank search box and the user's like, what do I do now? We, when you log into Logly, you actually see Some charts and graphs that help…
AI assessment note: “there's really, um, you know, I would say two themes that we're interested in”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q key aspects here, and you alluded to that a little bit, um, in an enterprise context, so, Ideally, you would want to get the corporate data, but also get external data. How do you, which obviously is going to make, you know, certain people scream within corporations. How do you, how do you think about that? And if you cannot get the corporate data, how do you train the algorithms?
A That's a great question. So, the data is a huge piece of it, right? It's garbage in, garbage out, as, as most of you all know in general data science, right? And so, um, there are actually a broad variety of Things, ways that we deal with this. For some things, for instance, we're collecting our own data. We're paying for it. We're deduplicating, cleaning, labeling it ourselves. So, for instance, it sounded like a gimmick in the beginning, but a lot of people are asking us to do food classification for General sort of obesity, diabetes kinds of applications, fitness, and so on. Uh, and so we collected our own food data set with hundreds of classes, uh, and basically just sell the classifier as is. But then there are other data sets that we will never get, and, uh, we won't ever get in the future unless we re-change, which is medical, for instance. So, uh, one of our biggest partners, VRAD Virtual Radiologic, who are the largest teleradiology provider in the United States, And they have an amazing treasure trove of, of data, for instance, for intracranial hemorrhage, so classifying brain bleeds in, uh, three volumes or CT scans. And so, with them, we really have to partner, and, and then the ideal scenario, which is actually the case, uh, for VRED, we're then allowed to also sell those classifiers afterwards. But the most common case is companies just have their own data set, th…
AI assessment note: “there are actually a broad variety of Things, ways that we deal with this”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Great, thank you so much. Um, I'll, I'll pass on to my two people here, but, uh, what does a data science team look like, or who, who does all of this?
A Um, so, a data science team, in our case, it's, um, Uh, six data scientists, um, as well as, uh, several data science, uh, what we call data science engineers, and so on the engineering side, uh, those would be folks who would be moving data around and creating all this infrastructure, and data science would be the guys who are actually, uh, modeling for things, uh, developing personalization engines, uh, trying to find interesting, uh, tidbits of information inside data, Um, and work, ultimately working with the business. So one thing which is, um, we try to be very, um, very hard about is that whatever we do, it ultimately needs to be solving some kind of a problem in a business. Problem in a business doesn't mean necessarily for OpenTable. Problem for our customers, for diners, for restaurants. Um, so we look at a more pragmatic approach to data science, and, and, um, that's on one side. Lastly, data science team would be working on things like inventory optimization inside restaurants. So that's also, I didn't touch on that area at all, but It's a big area of how the tables and slots of tables should be allocated. You know, if a restaurant opened from six to nine and somebody takes a seven p.m. reservation, they essentially block that restaurant for the whole night, rather than somebody would take a six o'clock or eight o'clock, and then a restaurant can fit two parties tha…
AI assessment note: “six data scientists, um, as well as, uh, several data science”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q And so the, the idea of the network is that some of this is shared? So some of the data is shared for research purposes?
A Yeah, so if you, if you work with Flatiron as a provider, so if you are a cancer center, um, one of the things you agree, in addition to kind of sending us your complete copy of your electronic health record, all the documents, we, we see everything, um, on a nightly basis, you agree to allow us to process and de-identify Your data so that we can aggregate it amongst the broader cohort. And in return, we are kind of providing folks with value props from the network itself. So you could think of things like benchmarking or the ability to do research on broader cohorts that any one single cancer center might not be able to do. They only see, you know, 3000 patients a year. We see a little under 700,000. So there are, Fred Wilson writes the best kind of network build out. Which I really, I really love. He talks about the single player use case and the multiplayer use case, and that's what we have. We have a, a single player use case, which is, we help you use your data better, and then we have a multiplayer use case, which is, oh no, by the way, now we've got this network behind us. Here are values that you can derive from the fact that there's, you know, 2000 other doctors also sharing their data. And that's the, the single player, multiplayer, um, game.
AI assessment note: “you agree to allow us to process and de-identify Your data so that we can aggregate it”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q So I assume you have to, you know, a little bit to Mark's presentation in a slightly different field, although I guess there is a big genomics component to all the cancer data, but, um, so do you see a lot of stuff like that, like the, just like crazy dirty data in lots of different formats?
A Yeah, I would say, we talk about all the, the fun software, but the reality is we spend most of our time cleaning data. Um, it's like very non-glamorous, uh, kind of work and infrastructure. Um, so this is real-world data. This is a, this is the equivalent of, like, what the doctor used to write by hand in a chart, but just now typed in a chart. Um, but the amount of kind of structured, normalized data that we see at the source is extraordinarily limited. Um, we see lab values with different units of measurement. We see sometimes scientific names for drugs versus generic names versus, you know, um, brand names, for example. Uh, so on the structured data side, there's this huge, what I would call, kind of, data normalization effort. Um, that we have behind the scenes, so we, we map. Um, much of this is actually done with people. Uh, and then we have this really unique infrastructure around unstructured data, which is all the notes, the pathology reports, where the genomic data actually ends up showing up. Uh, you know, you get a report, and when, when I say get a report, what, what that typically means is the report is faxed to you as the physician, and you, you read it, and then you, you hand it to your, your, uh, medical records person, and they scan it in. And so we see those. They're, you know, you can tell they're tilted, and sometimes there's like a crease down the middle,…
AI assessment note: “the reality is we spend most of our time cleaning data.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q this, and thanks for, um, not turning this into a product pitch, which is something I typically ask people to do, but they don't always do that. Um, could, but so, since you've been so kind from that, um, perspective, could you actually, uh, tell us a little bit more about what, uh, Recommend actually does, just, you know, in a, in a minute, the, the first product in particular?
A Absolutely. So I mentioned single gene disorders like cystic fibrosis. Those are actually regularly tested for in the U.S. by either fertility physician, reproductive endocrinologists, or OBGYNs. Um, LabCorp and Quest, two of the two biggest labs, dominate that testing market, um, and each disease that you might test for, cystic fibrosis, spinal muscular atrophy, Fragile X, Tay-Sachs, these are well-known diseases, can cost thousands of dollars using older technologies, so the first thing we did actually is apply these microarray and sequencing technologies, and we developed a test that can do 250 of the most common diseases of this nature for under 500 dollars to patients that don't have private insurance, and This is just making it more accessible. It's generating more data. It's more accurate. It's faster. It's really an application of these technologies to an area that we're already regularly doing in medicine, um, and beyond that, what we can do now is we can leverage the amount of data we're generating at the point of care and actually run research and development and become an asset for our physicians, so, um.
AI assessment note: “we developed a test that can do 250 of the most common diseases”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q So, so what's Hell's tap now? What, what's, what does he do?
A Sure, HealthUp is, is very simple. It's, it's, ah, the first ever, ah, end-to-end experience in, in healthcare. So from the moment you're not feeling great, all the way to the moment you're feeling good, you can use HealthUp, like, as a health utility, right? You come to HealthUp, you have a question. We have a network with more than 66,000 physicians, ah, U.S. licensed physicians in good standing, and you can ask any health question and get an answer from a doctor in seconds or in minutes. Ah, you can use a repository. We serve more than 2.6 billion doctor answers to date. So it's a very extensive repository, and a lot of data about what people actually find valuable, but not only what doctors answered to people before, but actually doctors peer review each other for quality. Right, so every answer that is given to people in HealthTap goes back to the doctors that either agree with the answer, or if they disagree, they add a comment, Or they add their own answers, so the database is the world's most extensive database of doctor knowledge that is organized by how people ask questions. And beyond that, as we start getting a lot of engagement, we start getting more and more doctors, which basically gave them the opportunity to create tips, ah, to review news online, and as of recent, to actually rate apps. You know, there are more than a 100,000 health apps on, ah, on Google Play…
AI assessment note: “you can ask any health question and get an answer from a doctor in seconds”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q so, um, you basically gather a lot of data about every person, right? So help us understand, so this is the world of EMRs, that has your history, and that's the history that you have with your doctor, this is yet another history. What, how do you think about this, and how do you protect it, how do you, and what things do you do once you have my data?
A Yeah, it's pretty cool. I mean, like, I think that we try to keep it in context. So first of all, patients on HealthTap are always anonymous, Right? So, you never see who the patient is. When you ask it, ah, doctors are not. Doctors are very visible, right? You see who the doctor is answering the question, their credential. Each and every one of them has a virtual practice. You can see where they practice. You can contact them, so it's very transparent. On the patient side, nothing. You cannot see any other patients on HealthTap. It's all completely private and encrypted. Ah, the other thing that is very important is that we are doing things in context. So when you ask a question on HealthTap, we will ask you to add Uh, three attributes. We're using our knowledge base and some machine learning techniques to actually try to attach to your question some related data points that will help the doctor give you a more personalized answer. And why is that important? Because if, ah, a 26 year old woman with, ah, that, that is pregnant is asking what are the potential implications of diabetes and she has no, no, ah, other conditions, she's taking no medication, and the same question exactly is asked by a 76 year old guy with three comorbidities, Taking four medications, the same semantic question will get a very different answer. Right, so adding these attributes to the question and tag…
AI assessment note: “It's all completely private and encrypted.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q So is this why you needed a parallel processing to be able to do that? Because you essentially have multiple layers operating at the same time, or no?
A Right. So the problem with this is that If you want the network to be able to recognize images properly or speech, you need them to be very, very large on the order of, um, so essentially the elementary operations that each of those, uh, elements are doing, uh, is, is, you know, multiply and, you know, multiplication by number and addition. And you may have, you know, in a typical conventional net, you may have something like between one and ten billion, uh, multiply, accumulate, operations, where each multiplication is a coefficient subject to learning. Ok. So you have a very large system. Computing the output takes, you know, five, ten billion operations. And you can't do this on the CPU. It's just too slow. So you have to go to GPUs. Current GPU cards are capable of, you know, four or five teraflops for a single GPU. You can parallelize on multiple GPUs, like something like four or eight in a single machine. And the problem, of course, is that you have to do this millions of times because you need to train the The network on millions of images before it's able to do any kind of proper recognition, and you have to cycle through those images perhaps a hundred times before it gets it right. So, uh, you know, the first kind of such systems, of course, convolutional nets have been, have been around for a long time, but the first large convolutional nets that are appropriate for o…
AI assessment note: “Right. So the problem with this is that If you want the network”