Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered produced feed
D 5 · C 5 · P 5 · Cm 4 4.85
Q I'm so pleased you said about supervised learning there, because we've also spoken before about the derivative data via supervised learning and comparing that to initial data. So maybe for those that don't know, how do the two compare? Let's start with that.
A Sure. So a great example is in the case, we're coming back to the case of Domisto. The original data is what did this human being do to resolve an incident? And I might make the recommendation to you to say, hey, Harry, they did these three steps. Hit this green button or make this rule. And that's the original data, and that's great, and that can take you very far, and for instance, in the case of Google PageRank, it took them quite far in the early days. The derivative data that you can, you know, perform supervised learning on is really this feedback loop. So in the Domisto case, it's your ability through further interactions with the product to say, hey, Domisto, you actually got this wrong. I'm going to interact with the system a little bit more and teach you some further nuance that you can then learn on and do a better job next time. Just as Google takes the clicks on, you know, which URL you actually selected was the most interesting and feeds that back into their search algorithm results to the point where that's now a huge part of how they give you the most relevant piece of content.
AI assessment note: “The original data is what did this human being do... The derivative data... is really this feedback loop.”
Answered produced feed
D 5 · C 5 · P 5 · Cm 4 4.85
Q So latching onto that core thesis around the importance of data capture, you previously espoused, to me, a cruel walk-run approach in terms of regards to data. What do you mean by this, and how does that affect then your thinking when assessing the data set an opportunity presents?
A Back when I worked at Splunk, people didn't throw around the word AI at the time, but they would say, hey, Jake, how can I get some predictive analytics for X? And I would say, you know, we can talk about that, But can you tell me what's going on in your organization today? Do you have a live insight feed into what's going on in your organization? Or do you have a historical insight feed for the last two weeks? And oftentimes, the answer to those questions was no. And so part of what I mean, like crawl, walk, run, is let's sort of understand the past and the present before we start thinking about the future. But more practically speaking, it comes back to this notion of workflow first, data capture, followed by model building and iterations. And maybe we can look through the lens of one of my portfolio companies called Domisto. This is in the security space. You can think of Domisto as a command line on steroids for security that plugs in and integrates with every single security tool that an Build an awful lot of workflow and help and tooling to make this an enjoyable experience, or the incident responders will just go back to their security command line or whatever tools are currently in joy. But once they build a good enough workflow experience to get the user to stay in the tool, they're able to harness this really interesting data set that no one has centralized for the fi…
AI assessment note: “understand the past and the present before we start thinking about the future”
Answered produced feed
D 5 · C 5 · P 5 · Cm 4 4.85
Q No, that makes sense. In terms of the data sets themselves, I actually had Aaron Fandavender from Founders Fund on the show the other day, and he said the value of large data sets is mostly overplayed. I'm intrigued. Would you agree with this? And how do you analyze those inflection points whereby increases in data really don't amount to increases in value derived?
A Yeah, I think I'll, I'll have to cop out and sort of say it depends, right? I think there's certain situations in which the value of large data is not overplayed. Certainly there's some resonance with what Aaron says. And so far as if I've collected a petabyte of network traffic data, the incremental one data point is inherently not going to be valuable to me. The question is, how defensible is that increment of data, and then how valuable is that net increment? And so, for instance, if you were doing something in the medical space, If you had somehow locked up a contractual relationship to the largest repository of these images, that's very, very valuable. Particularly if the rest of the market is fragmented, it would make it very hard for the next player to amalgamate or amass equal amounts of data. Now, if they had twenty million images, is the 20,000,001st image going to be that valuable? No. But that's where I think a lot of this feedback loop and sort of supervised learning around the data is Actually becomes a little bit more of an advantage than the original data set itself.
AI assessment note: “I think there's certain situations in which the value of large data is not overplayed.”
Answered produced feed
D 5 · C 5 · P 5 · Cm 4 4.85
Q Can I ask, how do you view automated data creation? We often see it with self-driving cars where they test in simulated environments where the data itself is actually self-created. How do you think about those sorts of environments where the data is automated in terms of creation?
A So, you know, I think there's a, there's a whole body of artificial intelligence called reinforcement learning or trial-based learning, and there's analogies for even sort of just the data creation, and I think it's, it's sort of a wonderful way to train algorithms for a certain set of tasks, and so in the instance of playing video games, it makes all the sense in the world, and you can train these computer systems to do things that no human would have ever thought to do before. In the instance of self-driving cars, It could make a lot of sense, and I think it can get you 99.9% of the way there, but the question is, how fully reflective is the underlying data set of the real world, and in the instance where it's not 100% fully reflected, can you actually generate all possible edge cases, and are you willing to accept that level of risk if there's point oh one percent chance that they didn't get this right? It's very, very different when we're playing a video game versus when a human life may actually be at stake.
AI assessment note: “In the instance of self-driving cars, It could make a lot of sense”
Answered produced feed
D 5 · C 5 · P 4 · Cm 4 4.60
Q I do want to touch on an element you said earlier, though, being kind of the exclusivity of data and having that sole relationship to that data, making it your mode. I'm intrigued with the environment today with GAFA, Google, Amazon, Facebook, and Apple. Are you concerned in terms of startups accessibility, especially with the incumbency data advantage that exists today?
A Uh, yes and no. It's something I think about a lot, to the extent that a large internet company like a Google has, one, both the data set, and two, Really the desire to use that data set and to monetize it. So for instance, if you want to use some Google-like or Google-specific dataset for an advertising product, it's a really uncomfortable situation because at any point in time, Google may very well decide to build that same product. Now, for instance, in security, could Google build some tremendously interesting datasets based on all of their internals and operating procedures? Sure. Is that in their mission? Are they likely to do that? No. So I think it very much depends exactly what the product is and if there's both Data at these internet companies and intent. And if you believe there could be both of those, that's going to be a tricky situation for a long time to come.
AI assessment note: “it very much depends exactly what the product is and if there's both Data at these internet companies and intent”
Answered produced feed
D 5 · C 5 · P 4 · Cm 4 4.60
Q I'm so pleased you said about supervised learning there, because we've also spoken before about the derivative data via supervised learning and comparing that to initial data. So maybe for those that don't know, how do the two compare? Let's start with that.
A Sure. So a great example is in the case, we're coming back to the case of Domisto. The original data is what did this human being do to resolve an incident? And I might make the recommendation to you to say, hey, Harry, they did these three steps. Hit this green button or make this rule. And that's the original data, and that's great, and that can take you very far, and for instance, in the case of Google PageRank, it took them quite far in the early days. The derivative data that you can, you know, perform supervised learning on is really this feedback loop. So in the Domisto case, it's your ability through further interactions with the product to say, hey, Domisto, you actually got this wrong. I'm going to interact with the system a little bit more and teach you some further nuance that you can then learn on and do a better job next time. Just as Google takes the clicks on, you know, which URL you actually selected was the most interesting and feeds that back into their search algorithm results to the point where that's now a huge part of how they give you the most relevant piece of content.
AI assessment note: “The original data is what did this human being do to resolve an incident”
Answered produced feed
D 5 · C 5 · P 4 · Cm 4 4.60
Q So latching onto that core thesis around the importance of data capture, you previously espoused, to me, a cruel walk-run approach in terms of regards to data. What do you mean by this, and how does that affect then your thinking when assessing the data set an opportunity presents?
A Back when I worked at Splunk, people didn't throw around the word AI at the time, but they would say, hey, Jake, how can I get some predictive analytics for X? And I would say, you know, we can talk about that, But can you tell me what's going on in your organization today? Do you have a live insight feed into what's going on in your organization? Or do you have a historical insight feed for the last two weeks? And oftentimes, the answer to those questions was no. And so part of what I mean, like crawl, walk, run, is let's sort of understand the past and the present before we start thinking about the future. But more practically speaking, it comes back to this notion of workflow first, data capture, followed by model building and iterations. And maybe we can look through the lens of one of my portfolio companies called Domisto. This is in the security space. You can think of Domisto as a command line on steroids for security that plugs in and integrates with every single security tool that an Build an awful lot of workflow and help and tooling to make this an enjoyable experience, or the incident responders will just go back to their security command line or whatever tools are currently in joy. But once they build a good enough workflow experience to get the user to stay in the tool, they're able to harness this really interesting data set that no one has centralized for the fi…
AI assessment note: “crawl, walk, run, is let's sort of understand the past and the present”
Answered produced feed
D 5 · C 5 · P 4 · Cm 4 4.60
Q Can I ask, how do you view automated data creation? We often see it with self-driving cars where they test in simulated environments where the data itself is actually self-created. How do you think about those sorts of environments where the data is automated in terms of creation?
A So, you know, I think there's a, there's a whole body of artificial intelligence called reinforcement learning or trial-based learning, and there's analogies for even sort of just the data creation, and I think it's, it's sort of a wonderful way to train algorithms for a certain set of tasks, and so in the instance of playing video games, it makes all the sense in the world, and you can train these computer systems to do things that no human would have ever thought to do before. In the instance of self-driving cars, It could make a lot of sense, and I think it can get you 99.9% of the way there, but the question is, how fully reflective is the underlying data set of the real world, and in the instance where it's not 100% fully reflected, can you actually generate all possible edge cases, and are you willing to accept that level of risk if there's point oh one percent chance that they didn't get this right? It's very, very different when we're playing a video game versus when a human life may actually be at stake.
AI assessment note: “it's sort of a wonderful way to train algorithms for a certain set of tasks”
Answered produced feed
D 5 · C 5 · P 4 · Cm 4 4.60
Q No, that makes sense. In terms of the data sets themselves, I actually had Aaron Fandavender from Founders Fund on the show the other day, and he said the value of large data sets is mostly overplayed. I'm intrigued. Would you agree with this? And how do you analyze those inflection points whereby increases in data really don't amount to increases in value derived?
A Yeah, I think I'll, I'll have to cop out and sort of say it depends, right? I think there's certain situations in which the value of large data is not overplayed. Certainly there's some resonance with what Aaron says. And so far as if I've collected a petabyte of network traffic data, the incremental one data point is inherently not going to be valuable to me. The question is, how defensible is that increment of data, and then how valuable is that net increment? And so, for instance, if you were doing something in the medical space, If you had somehow locked up a contractual relationship to the largest repository of these images, that's very, very valuable. Particularly if the rest of the market is fragmented, it would make it very hard for the next player to amalgamate or amass equal amounts of data. Now, if they had twenty million images, is the 20,000,001st image going to be that valuable? No. But that's where I think a lot of this feedback loop and sort of supervised learning around the data is Actually becomes a little bit more of an advantage than the original data set itself.
AI assessment note: “I'll have to cop out and sort of say it depends”
Answered produced feed
D 5 · C 5 · P 4 · Cm 4 4.60
Q I love that as an entry point, but I do have to ask, what were the big takeaways for you from spanning very different organizations from Lockheed Martin to Splunk and Cloudera? What were the big takeaways for you from those times in operations?
A Well, look, I think there's an inherent inefficiency of scale at any company, and it's no disrespect to Lockheed Martin. Once an organization reaches a certain size, the amount of overhead and And planning makes it very, very hard to move nimbly, and quite frankly, the incumbent nature of your business can hold you back. We talk about this in the classic Clay Christensen theory of disruption. Do you actually have what it takes to move quickly and disrupt yourself? And some of the things I observed at Lockheed have been very informative in terms of how I work with my startups to help them build their organizations and continually ask that question of, Hey, there's no sacred cows. If we need to go back to the drawing board, if we need to give up this line of revenue stream for something that's going to be more exciting in the future, I think that gets harder and harder for you to do as the company grows. You really need to instill that value set in companies when they're young.
AI assessment note: “I think there's an inherent inefficiency of scale at any company”
Answered produced feed
D 5 · C 5 · P 4 · Cm 4 4.60
Q I'm intrigued. From a startup perspective, how do you think about accessibility of data sets and their ability to, if they identify the workflow, then actually capture that data in an effective manner? And how have you seen that play out for startups today?
A Yeah, so I think the real question is this cold start T equals zero problem. If you come up to me and say, hey, I have this great Whatever it is, vertical SaaS solution for Model X, just let me hook up to your system and wait six months and I'll collect a lot of data. You know, that's a tough sales motion. If you have some way to bootstrap that problem by partnering or somehow acquiring a unique or defensible, you know, maybe with a legal agreement set of data that no one else can get their hands on, that's a great strategy for bootstrapping. Or it falls back to this notion of workflow where I do actually provide an exciting amount of value even if I don't have the data yet. And that's really where I think the bulk of the opportunity is going to be going forward, particularly in areas like vertical SaaS, where I can improve the current experience without any data, and any data that I can collect to improve your life is just gravy on top.
AI assessment note: “falls back to this notion of workflow where I do actually provide an exciting amount”
Answered produced feed
D 5 · C 5 · P 4 · Cm 4 4.60
Q I do want to touch on an element you said earlier, though, being kind of the exclusivity of data and having that sole relationship to that data, making it your mode. I'm intrigued with the environment today with GAFA, Google, Amazon, Facebook, and Apple. Are you concerned in terms of startups accessibility, especially with the incumbency data advantage that exists today?
A Uh, yes and no. It's something I think about a lot, to the extent that a large internet company like a Google has, one, both the data set, and two, Really the desire to use that data set and to monetize it. So for instance, if you want to use some Google-like or Google-specific dataset for an advertising product, it's a really uncomfortable situation because at any point in time, Google may very well decide to build that same product. Now, for instance, in security, could Google build some tremendously interesting datasets based on all of their internals and operating procedures? Sure. Is that in their mission? Are they likely to do that? No. So I think it very much depends exactly what the product is and if there's both Data at these internet companies and intent. And if you believe there could be both of those, that's going to be a tricky situation for a long time to come.
AI assessment note: “Uh, yes and no. It's something I think about a lot”
Answered produced feed
D 4 · C 5 · P 5 · Cm 4 4.55
Q Not at all, but I'd love to start today with a little on you and how you made your foray into the world of early stage investing. What was your entry point?
A Sure. So I'll give you a quick rundown on, on my background. Undergrad, I went to Duke and I did electrical engineering, computer science, and economics. And that was my sort of way of saying that things like photography that had what I believe to be subjective judgment involved were, were, were scary to me. Afterwards, I went to Lockheed Martin through a great engineering rotational program that taught me a lot of what's right and wrong about how large companies operated. After a couple years of that, though, I got a little tired of it, and I said, hey, I want to get away from 150,000 people, and I want to see what something smaller is like. And so I wound up going to business school, and at business school, I came across these two trends. As I started reading cloud and data, and the more and more I read about data and cloud, I realized in 2008, you know, cloud was sort of being taken care of. There was this thing called Amazon Web Services that was taking off, but on the other hand, data, the more I read about it, the more I just felt like, hey, I actually know a little bit about this. Hey, I know more than most people about this. Why don't I try to do something in the data space? And so like any good first year MBA student, I made a list of 20 companies, cold called the CEOs, and wound up at Cloudera in the very early days, when it was just the four founders, eight engineers…
AI assessment note: “So I'll give you a quick rundown on, on my background.”
Answered produced feed
D 4 · C 5 · P 4 · Cm 4 4.30
Q I love that as an entry point, but I do have to ask, what were the big takeaways for you from spanning very different organizations from Lockheed Martin to Splunk and Cloudera? What were the big takeaways for you from those times in operations?
A Well, look, I think there's an inherent inefficiency of scale at any company, and it's no disrespect to Lockheed Martin. Once an organization reaches a certain size, the amount of overhead and And planning makes it very, very hard to move nimbly, and quite frankly, the incumbent nature of your business can hold you back. We talk about this in the classic Clay Christensen theory of disruption. Do you actually have what it takes to move quickly and disrupt yourself? And some of the things I observed at Lockheed have been very informative in terms of how I work with my startups to help them build their organizations and continually ask that question of, Hey, there's no sacred cows. If we need to go back to the drawing board, if we need to give up this line of revenue stream for something that's going to be more exciting in the future, I think that gets harder and harder for you to do as the company grows. You really need to instill that value set in companies when they're young.
AI assessment note: “there's an inherent inefficiency of scale at any company”
Answered produced feed
D 4 · C 5 · P 4 · Cm 4 4.30
Q I'm intrigued. From a startup perspective, how do you think about accessibility of data sets and their ability to, if they identify the workflow, then actually capture that data in an effective manner? And how have you seen that play out for startups today?
A Yeah, so I think the real question is this cold start T equals zero problem. If you come up to me and say, hey, I have this great Whatever it is, vertical SaaS solution for Model X, just let me hook up to your system and wait six months and I'll collect a lot of data. You know, that's a tough sales motion. If you have some way to bootstrap that problem by partnering or somehow acquiring a unique or defensible, you know, maybe with a legal agreement set of data that no one else can get their hands on, that's a great strategy for bootstrapping. Or it falls back to this notion of workflow where I do actually provide an exciting amount of value even if I don't have the data yet. And that's really where I think the bulk of the opportunity is going to be going forward, particularly in areas like vertical SaaS, where I can improve the current experience without any data, and any data that I can collect to improve your life is just gravy on top.
AI assessment note: “provide an exciting amount of value even if I don't have the data yet”
Partly produced feed
D 3 · C 5 · P 4 · Cm 4 4.00
Q Not at all, but I'd love to start today with a little on you and how you made your foray into the world of early stage investing. What was your entry point?
A Sure. So I'll give you a quick rundown on, on my background. Undergrad, I went to Duke and I did electrical engineering, computer science, and economics. And that was my sort of way of saying that things like photography that had what I believe to be subjective judgment involved were, were, were scary to me. Afterwards, I went to Lockheed Martin through a great engineering rotational program that taught me a lot of what's right and wrong about how large companies operated. After a couple years of that, though, I got a little tired of it, and I said, hey, I want to get away from 150,000 people, and I want to see what something smaller is like. And so I wound up going to business school, and at business school, I came across these two trends. As I started reading cloud and data, and the more and more I read about data and cloud, I realized in 2008, you know, cloud was sort of being taken care of. There was this thing called Amazon Web Services that was taking off, but on the other hand, data, the more I read about it, the more I just felt like, hey, I actually know a little bit about this. Hey, I know more than most people about this. Why don't I try to do something in the data space? And so like any good first year MBA student, I made a list of 20 companies, cold called the CEOs, and wound up at Cloudera in the very early days, when it was just the four founders, eight engineers…
AI assessment note: “So I'll give you a quick rundown on, on my background.”
Redirected produced feed
D 3 · C 4 · P 4 · Cm 4 3.70
Q I do want to then delve on two threats, because we spoke about those two threats. What are some other potential threats that you particularly identify in the market, and where do you see those inherent gaps?
A Yeah, so I think there's a few areas of security that are quite interesting these days. I think there's also a few that are overplayed. Getting back to our conversation around AI and ML, one area that I think is unfortunately slightly overplayed is this machine learning anomaly detection to rule them all off by itself. And the train of thought goes, hey, I can help you. I can do this advanced analytics, help you identify something anomalous in your organization, and it's worth investigating. And that actually turns out to be great. And in fact, here's the recipe or playbook how to resolve the problem that you can give to the twenty-one-year-old that just joined my security team two weeks ago. Some of these CISOs might say, hey, I don't even want to know about this anomaly because there's nothing I can do about it. So we've over-rotated a little bit too much on the ML side. There's a few areas that I'm tremendously excited about. The first is this notion of security orchestration and validation. What I mean by that is a lot of CISOs have already made a tremendous number of investments in security products. And the 287th security product layer X for a new threat vector Y, maybe they'll tolerate that. It forced to, but they're, they're kind of sick of that. And rather, they would much like to have this conversation of, hey, how can I get more out of what I've already got? And whet…
AI assessment note: “there's a few areas of security that are quite interesting these days.”
Partly produced feed
D 3 · C 4 · P 4 · Cm 4 3.70
Q I do want to then delve on two threats, because we spoke about those two threats. What are some other potential threats that you particularly identify in the market, and where do you see those inherent gaps?
A Yeah, so I think there's a few areas of security that are quite interesting these days. I think there's also a few that are overplayed. Getting back to our conversation around AI and ML, one area that I think is unfortunately slightly overplayed is this machine learning anomaly detection to rule them all off by itself. And the train of thought goes, hey, I can help you. I can do this advanced analytics, help you identify something anomalous in your organization, and it's worth investigating. And that actually turns out to be great. And in fact, here's the recipe or playbook how to resolve the problem that you can give to the twenty-one-year-old that just joined my security team two weeks ago. Some of these CISOs might say, hey, I don't even want to know about this anomaly because there's nothing I can do about it. So we've over-rotated a little bit too much on the ML side. There's a few areas that I'm tremendously excited about. The first is this notion of security orchestration and validation. What I mean by that is a lot of CISOs have already made a tremendous number of investments in security products. And the 287th security product layer X for a new threat vector Y, maybe they'll tolerate that. It forced to, but they're, they're kind of sick of that. And rather, they would much like to have this conversation of, hey, how can I get more out of what I've already got? And whet…
AI assessment note: “The first is this notion of security orchestration and validation.”