Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q it and get better and better at it. It's a really important start, I think. I have to ask, Scale is a generational defining company in many respects. Just raised it at twenty-five billion dollars. It ends this year at two billion dollars. Fucking nuts. Well done, Alex. You were there for, like, four and a half years. How did your time at scale shape how you think about product?
A Yeah, I think maybe a few things to touch on. So one, I think, you know, scale was at the frontier of a lot of different AI movements over the last, like, five, six years, and I think one thing we did really well was listening to frontier customers, customers that are at the frontier of whatever we were pushing, and because we were Just at the cusp of new markets every single time, um, it was super beneficial to get deep with those customers because you could have conviction that they would basically define the market and everything they wanted, you know, everyone else would follow. And so, you know, early on in self-driving, uh, Neuro was super, super early on in a lot of LIDAR and a lot of really hard, what's called pedestrian work, uh, like pedestrian labeling work. And so we ended up doing a lot of kind of custom things for Neuro, getting super deep with them. Lo and behold, everyone else started asking for the same thing. I think whenever you're at these cutting edge markets, and you know, today, today's world is definitely the case, find the customers who are really good at what they do, and then, you know, chase them, even over a fit to them, um, as long as you have conviction that, you know, everyone else will follow.
AI assessment note: “find the customers who are really good at what they do, and then, you know, chase them”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q it and get better and better at it. It's a really important start, I think. I have to ask, Scale is a generational defining company in many respects. Just raised it at twenty-five billion dollars. It ends this year at two billion dollars. Fucking nuts. Well done, Alex. You were there for, like, four and a half years. How did your time at scale shape how you think about product?
A Yeah, I think maybe a few things to touch on. So one, I think, you know, scale was at the frontier of a lot of different AI movements over the last, like, five, six years, and I think one thing we did really well was listening to frontier customers, customers that are at the frontier of whatever we were pushing, and because we were Just at the cusp of new markets every single time, um, it was super beneficial to get deep with those customers because you could have conviction that they would basically define the market and everything they wanted, you know, everyone else would follow. And so, you know, early on in self-driving, uh, Neuro was super, super early on in a lot of LIDAR and a lot of really hard, what's called pedestrian work, uh, like pedestrian labeling work. And so we ended up doing a lot of kind of custom things for Neuro, getting super deep with them. Lo and behold, everyone else started asking for the same thing. I think whenever you're at these cutting edge markets, and you know, today, today's world is definitely the case, find the customers who are really good at what they do, and then, you know, chase them, even over a fit to them, um, as long as you have conviction that, you know, everyone else will follow.
AI assessment note: “find the customers who are really good at what they do, and then, you know, chase them”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q What have been the biggest mistakes you made in user testing, or biggest mistakes you see other people make in user testing?
A I think one of them is biasing the users for what you want them to think. This is most useful, maybe with an example. So we built this, um, product called vault last June-ish and the essential thing is it allows you to do large scale extraction of documents. Like if you have a bunch of lease agreements, credit agreements, and you want 10, 20, 50 deal points out of it, uh, out of each one, it creates a whole grid for you. And the way we went out building that is when you upload a bunch of files, You say, Hey, I want these 10 things and Harvey will decompose your initial prompt and suggest like, okay, here's the 10 things. Am I on the right track and, you know, check, check, check. And then you launch it. It demos really well, um, because it's like takes your prompt and then converts it and shows a little thing. And when we showed it to people, people were like, oh my God, wow, that's amazing. And then we didn't really put it into the hands of people for live matters, live, live use cases. We just showed it and. They were like, oh my God, cool. It's a demo. And it's like, you know, interpreting my intent. The thing that we did wrong there was actually people want to just make those terms in the grid itself. People don't want the AI to, to convert it. They just want to say for these 10,000 agreements, I just want this one thing created in the grid itself. They want to get really t…
AI assessment note: “I think one of them is biasing the users for what you want them to think.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Can I ask, why is Claude better than OpenAI?
A Claude, so we did this, we did these evals recently. Claude three seven in particular, for legal reasoning, it's, it's, it's better at long form legal reasoning and drafting long form. Outputs. So a whole section of a merger agreement, for example, and making it super consistent, it's really good for. And then things like, um, extraction where extracting some terms form SPA, like a share purchase agreement is very nuanced because the answer isn't a verbatim text in the, in the agreement. It's like, it's getting a little in the weeds on legal, but if you're getting the indemnification cap on a share purchase agreement, You have to reason over four or five different clauses, because that cap is not a name thing in the deal. You have to, like, reason over the dependencies of four different clauses in that agreement, and then extract that out. For some of those types of extractions, Claude is, three seven in particular, is starting to get better. Every single model release that's come out from OpenAI, from one of the other competitors, um, we've benchmarked for since the beginning of time at Harvey, and three seven is really where we're starting to see some of Some of the performance better than OpenAI.
AI assessment note: “Claude three seven in particular, for legal reasoning, it's, it's, it's better at long form”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q What's a retro and what's a post-mortem?
A Yeah, so, so retros, retros are, you want to look back at some period of time and say, here's what we did well, here's what we didn't do well. And retros should be, you know, generally more regular, like every month, for example, at the end of the month, you say, What did we accomplish this month? What could have gone better? Here's what we should do for the next month. Postmortems are when there's an incident or issue, or we made the wrong decision. You want to evaluate that very specific decision or event that happened. Like if there's an incident and our app goes down, you know, for X amount of minutes, you want to make sure you do a postmortem and see like why that happened. Well, how do we resolve it? How could we have prevented it? And so postmortems are more ad hoc. One thing we're starting to do more of and. This is again, me, me learning as a, as a product leader at, at Harvey is a premortems. Premortems are before you start something like a big project or big initiative, you want to sit down with your team and say, what does success look like? And what will prevent us from achieving that? And what will go wrong? Give me all your worries. I I'm generally a very anxious paranoid person when I'm operating. Uh, I assume everything's going to go wrong. And so maybe it's like self therapeutic for me to be like, Guys, I think this is gonna happen, it's gonna happen, it's gon…
AI assessment note: “retros are, you want to look back... Postmortems are when there's an incident”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q To what extent are your evals the same as the public evals that are done on benchmark? I had, um, Tamar, the CPO and president of Clean on the show recently, and she said actually that their evals were very, very different to the benchmark public displays of evals. Are yours correlated, or are yours wildly different?
A Yeah, so our evals are very different. A lot of public legal benchmarks, and then even things like, you know, scales, uh, humanities last exam, they're all multiple choice. I would love if legal work was multiple choice, but every, any lawyer will tell you there are like a million options of what you could do, and so one, it's like, Multiple choices is, is not the right thing. And so we created this, um, benchmark called Big Law Bench consists of tasks that are real billable work tasks that lawyers do on a daily basis at our biggest customers and big law firms. The nature of these tasks are they're, they're very much open-ended. The problem with open-ended tasks is like, how do you evaluate them consistently across tasks? Like if you're saying Harvey generate a chronology that is very different than You know, draft me a motion for summary judgment, which is, you know, another kind of like litigation, litigation tasks. Then what we had to do is create rubrics for each of those tasks as well. And those got very specific. And, um, basically we have like hundreds of different tasks and each task has a rubric now. And the benefit that we have, and, you know, we, we kind of architect from the beginning is we have a lot of lawyers that we have brought from, you know, big law. To work with our AI team, with our engineering product team who sit side by side with them and say, here's how…
AI assessment note: “Yeah, so our evals are very different. A lot of public legal benchmarks”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q What was the most uncomfortable thing you embraced?
A Um, so there's a few things in, in, Different parts of my life. Um, maybe in college, um, I willingly took probably, like, one of the hardest, like, courses at, at Carnegie Mellon, like, operating systems. You have to build operating systems from scratch, and I didn't have to take that. I could go to my degree without it. I had just heard a lot of crazy things of people staying up two nights in a row doing it, and I was just like, let's do it. Why not? Um, another thing is, uh, you know, out of college, like a lot of CMU, um, you know, graduates, uh, We were offered jobs at the big companies, Microsoft, Palantir, Stripe, whatever. I decided to not do that, move to LA, join my friend's company, which ended up failing in a year, but, uh, just something off the beaten path, and when else would I live in LA? And so I ended up doing that, um, and then meandered my way through, uh, through startups, instead of kind of going for the, for the big companies.
AI assessment note: “I willingly took probably, like, one of the hardest, like, courses at, at Carnegie Mellon”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Can I ask, why is Claude better than OpenAI?
A Claude, so we did this, we did these evals recently. Claude three seven in particular, for legal reasoning, it's, it's, it's better at long form legal reasoning and drafting long form. Outputs. So a whole section of a merger agreement, for example, and making it super consistent, it's really good for. And then things like, um, extraction where extracting some terms form SPA, like a share purchase agreement is very nuanced because the answer isn't a verbatim text in the, in the agreement. It's like, it's getting a little in the weeds on legal, but if you're getting the indemnification cap on a share purchase agreement, You have to reason over four or five different clauses, because that cap is not a name thing in the deal. You have to, like, reason over the dependencies of four different clauses in that agreement, and then extract that out. For some of those types of extractions, Claude is, three seven in particular, is starting to get better. Every single model release that's come out from OpenAI, from one of the other competitors, um, we've benchmarked for since the beginning of time at Harvey, and three seven is really where we're starting to see some of Some of the performance better than OpenAI.
AI assessment note: “Claude three seven in particular, for legal reasoning, it's, it's, it's better at long form”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q What have been the biggest mistakes you made in user testing, or biggest mistakes you see other people make in user testing?
A I think one of them is biasing the users for what you want them to think. This is most useful, maybe with an example. So we built this, um, product called vault last June-ish and the essential thing is it allows you to do large scale extraction of documents. Like if you have a bunch of lease agreements, credit agreements, and you want 10, 20, 50 deal points out of it, uh, out of each one, it creates a whole grid for you. And the way we went out building that is when you upload a bunch of files, You say, Hey, I want these 10 things and Harvey will decompose your initial prompt and suggest like, okay, here's the 10 things. Am I on the right track and, you know, check, check, check. And then you launch it. It demos really well, um, because it's like takes your prompt and then converts it and shows a little thing. And when we showed it to people, people were like, oh my God, wow, that's amazing. And then we didn't really put it into the hands of people for live matters, live, live use cases. We just showed it and. They were like, oh my God, cool. It's a demo. And it's like, you know, interpreting my intent. The thing that we did wrong there was actually people want to just make those terms in the grid itself. People don't want the AI to, to convert it. They just want to say for these 10,000 agreements, I just want this one thing created in the grid itself. They want to get really t…
AI assessment note: “I think one of them is biasing the users for what you want them to think.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q And they, when in a world where speed is everything, they allow for a much more efficient and speedy process. To what extent do you agree in the dictatorial nature of product decisions? Or actually, do you think brainstorming and discussion is important?
A So I believe in benevolent dictatorships. So Singapore, as an example, I totally agree. I think it's super important that you have very, you know, clear vision and like execution from the top, because otherwise you're going to end up a tragedy of commons. And so it is very clarifying to say, Hey, we are behind from this competitor do X, or we need to land these types of customers do Y. And that direction is extremely helpful from the founders. But I think you, especially as you're, you know, hyper growing, you end up with a lot of confusion or maybe the details don't actually end up in the way you want if you don't bring your team along the journey. And I think I've learned a lot of lessons from this, seeing scale, you know, at Harvey, we definitely have not nailed this, but.
AI assessment note: “So I believe in benevolent dictatorships.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q What would you most like to do in your role, but because of decision making or resources you're not able to do?
A Uh, go on podcasts. No, I'm already doing that. Um, I think, um, I would like to work on, so there's Harvey, like the productivity suite that we have today, which is the core product. And then there's Harvey that does a lot of the end to end work in collaboration with law firms and, and enterprises. And that is stuff like you, you know, that enterprises pay millions of dollars to do. Like, I want to just tackle that, that work. Um, that is way, you know, way out. You need to do research. You need UX, um, experimentation. Uh, so we'll get to that, but right now we have to focus on Harvey, the productivity suite, the software that lawyers use, and not the, the Harvey that, that does the work end to end.
AI assessment note: “right now we have to focus on Harvey, the productivity suite”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q you have distribution and you have product as president, You then get hypergrowth, and you get amazing customers adopt you, and you have this virality that you see within certain circles, and that's where I really like actually vertical software, because I think you get that virality much easier because of brands in those spaces. What are the first things to break as you move into hypergrowth in product teams?
A Yes, lots of things break. So maybe defining hypergrowth is every three to six months, let's say, your revenue is more than 1.5 x, 1.2 x, your Employee count is more than 1.5 to X. Like it's a different company every single quarter, let's say, or, you know, half. And so a few things that break are one, it becomes super hard to know what to actually prioritize. Like by definition, you're getting a lot of customers and a lot of demand and people are asking for a lot of different things. You don't know what direction to go. You know, this extends to engineers, it extends to the founders, PMs, whatever. You just don't know if you don't provide that clarity, then, um, you're going to do a lot of the wrong things. That's, I think, one. And then generally you also get either too many people making a decision or not enough people making a decision. Like you get a lot of kind of tragedy of the comments, like something's breaking and no one knows who should be responsible for it. Like, you know, for example, product enablement of sales, who is responsible for that? No one really thinks about that as much. Um, you know, when, when you're growing a sales team alongside a product team and Winston put me in charge of it. So now I'm figuring that out, but I think there, there's just a lot of things that break that Don't have owners, and so I think it's super important for the founder to provi…
AI assessment note: “a few things that break are one, it becomes super hard to know what to actually prioritize.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Is user choice good? If you think about the paradox of choice and users get confused when they have too much choice, if you're just told, I think humans are lazy, just tell them which one.
A On average, humans are lazy and just need to be told what's best for them. The whole model choice, and again, we, we fell into this trap. I think that whole paradigm was created for the, the, the Silicon Valley or like tech audience, because everyone loved playing around with like, oh, one versus GP four versus, you know, these models. We hadn't, I mean, we still haven't quite figured out What the models are good at such that you can route appropriately. The biggest problem with these models and we have this problem is just like the capabilities are very eval constrained. You don't know exactly which model is better for the job. Sometimes it's user preference. We do a lot of side by side testing, sometimes still not super clear. And so I think the whole industry as a whole kind of fell into this trap, put everything in a dropdown and give users a bunch of choice.
AI assessment note: “On average, humans are lazy and just need to be told what's best for them.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q you have distribution and you have product as president, You then get hypergrowth, and you get amazing customers adopt you, and you have this virality that you see within certain circles, and that's where I really like actually vertical software, because I think you get that virality much easier because of brands in those spaces. What are the first things to break as you move into hypergrowth in product teams?
A Yes, lots of things break. So maybe defining hypergrowth is every three to six months, let's say, your revenue is more than 1.5 x, 1.2 x, your Employee count is more than 1.5 to X. Like it's a different company every single quarter, let's say, or, you know, half. And so a few things that break are one, it becomes super hard to know what to actually prioritize. Like by definition, you're getting a lot of customers and a lot of demand and people are asking for a lot of different things. You don't know what direction to go. You know, this extends to engineers, it extends to the founders, PMs, whatever. You just don't know if you don't provide that clarity, then, um, you're going to do a lot of the wrong things. That's, I think, one. And then generally you also get either too many people making a decision or not enough people making a decision. Like you get a lot of kind of tragedy of the comments, like something's breaking and no one knows who should be responsible for it. Like, you know, for example, product enablement of sales, who is responsible for that? No one really thinks about that as much. Um, you know, when, when you're growing a sales team alongside a product team and Winston put me in charge of it. So now I'm figuring that out, but I think there, there's just a lot of things that break that Don't have owners, and so I think it's super important for the founder to provi…
AI assessment note: “a few things that break are one, it becomes super hard to know what to actually prioritize.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Do we lose the design stage in a world of seamless prototyping where anyone really with some competence can spin up a prototype pretty quickly and we don't have to spend a long time in the design pre-prototype phase?
A Short answer is no. I don't think you lose the design function or you lose, um, kind of designers. I think you actually get better designs because you can prototype much faster, especially with AI. I, I actually always say prototype before the PRD, whenever you have an idea, whether it's engineers, designers, PMs, whatever, you should do the prototype before, because you're not going to really know how to build the product or what the exact UX should be, particularly like, again, with AI, when these, these new patterns that we can establish without actually having, you know, you play with it, some internal customers play with that, you know, whatever it is. And then when you do the whole formal PRD kind of design process, whatever, You end up writing PRDs and whatnot, uh, and just, and doing proper designs, but because you've done all this pre-work and prototyping, you can actually save a lot of cycles, uh, and just move faster at the end of the day, and then ultimately get a better outcome.
AI assessment note: “Short answer is no. I don't think you lose the design function”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Is user choice good? If you think about the paradox of choice and users get confused when they have too much choice, if you're just told, I think humans are lazy, just tell them which one.
A On average, humans are lazy and just need to be told what's best for them. The whole model choice, and again, we, we fell into this trap. I think that whole paradigm was created for the, the, the Silicon Valley or like tech audience, because everyone loved playing around with like, oh, one versus GP four versus, you know, these models. We hadn't, I mean, we still haven't quite figured out What the models are good at such that you can route appropriately. The biggest problem with these models and we have this problem is just like the capabilities are very eval constrained. You don't know exactly which model is better for the job. Sometimes it's user preference. We do a lot of side by side testing, sometimes still not super clear. And so I think the whole industry as a whole kind of fell into this trap, put everything in a dropdown and give users a bunch of choice.
AI assessment note: “On average, humans are lazy and just need to be told what's best for them.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q And they, when in a world where speed is everything, they allow for a much more efficient and speedy process. To what extent do you agree in the dictatorial nature of product decisions? Or actually, do you think brainstorming and discussion is important?
A So I believe in benevolent dictatorships. So Singapore, as an example, I totally agree. I think it's super important that you have very, you know, clear vision and like execution from the top, because otherwise you're going to end up a tragedy of commons. And so it is very clarifying to say, Hey, we are behind from this competitor do X, or we need to land these types of customers do Y. And that direction is extremely helpful from the founders. But I think you, especially as you're, you know, hyper growing, you end up with a lot of confusion or maybe the details don't actually end up in the way you want if you don't bring your team along the journey. And I think I've learned a lot of lessons from this, seeing scale, you know, at Harvey, we definitely have not nailed this, but.
AI assessment note: “I believe in benevolent dictatorships. So Singapore, as an example, I totally agree.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q What would you most like to do in your role, but because of decision making or resources you're not able to do?
A Uh, go on podcasts. No, I'm already doing that. Um, I think, um, I would like to work on, so there's Harvey, like the productivity suite that we have today, which is the core product. And then there's Harvey that does a lot of the end to end work in collaboration with law firms and, and enterprises. And that is stuff like you, you know, that enterprises pay millions of dollars to do. Like, I want to just tackle that, that work. Um, that is way, you know, way out. You need to do research. You need UX, um, experimentation. Uh, so we'll get to that, but right now we have to focus on Harvey, the productivity suite, the software that lawyers use, and not the, the Harvey that, that does the work end to end.
AI assessment note: “I want to just tackle that, that work. Um, that is way, you know, way out.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q If I'm a PM, and I currently consider myself the glue, and I'm listening to you going, I want to be WD-Fordy, what can I do to be WD-Fordy, where I would have been glue?
A Yeah, I think this starts with coming to terms that you don't have to be the star of the show. I've made this mistake, I maybe sometimes still do. PMs have a main character syndrome sometimes, that they have to Be the CEO of the product, be this face of the product. The first thing you have to do is, is it actually the best thing for the company for me to do everything I do? It starts with that. It's maybe sounds simple. Maybe it sounds baseline. It's just give others the opportunity to shine for the skillset that they have. And then, you know, as I said, putting, putting designers, putting engineers in front of customers, giving them room to give ideas, you know, brainstorming with everyone in the room and a lot of kind of like tactical things, but it really starts with saying. I don't have to be the main character. I just have to make the product and the company succeed.
AI assessment note: “putting designers, putting engineers in front of customers, giving them room to give ideas”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q ChatGPT kind of made it the default by consumer expectations. Do you think that chat is the right interface?
A Definitely not. It's the command line starting point of this new frontier, like the MS-DOS was, you know, way back when. I think there's a few things in how, you know, how I think about it. I think One is chat is very linear and very one shot. You put in something and then you get an answer and then sure you can follow up and ask questions on top, but real work, you need to gather a lot of information to actually produce some work product, right? Um, you may have to get data from people that you don't even know have that data, um, or that context or whatever. And so, one principle that we use when building products is Something called the Ikea effect. The Ikea effect is like, you know, when, uh, Ikea started going super viral and super big, they made it super simple and nice and delightful to actually assemble the furniture. And what that did was build brand affinity of like, I put this together. I am responsible for it. You know, of course now people hire people to assemble things, but that, that just created this like cult, like following with Ikea. And so for us, how you can start to change this chat UX is like. How can, um, you know, the AI ask for more information or be like, hey, I wrote this first draft of something. Give me feedback on it before I continue. Like a real coworker, like a good coworker would do. So yeah, I think, I think chat is the very, very beginning, a…
AI assessment note: “Definitely not. It's the command line starting point of this new frontier”
Answered raw tape
D 5 · C 5 · P 4 · Cm 3 4.45
Q Do we lose the design stage in a world of seamless prototyping where anyone really with some competence can spin up a prototype pretty quickly and we don't have to spend a long time in the design pre-prototype phase?
A Short answer is no. I don't think you lose the design function or you lose, um, kind of designers. I think you actually get better designs because you can prototype much faster, especially with AI. I, I actually always say prototype before the PRD, whenever you have an idea, whether it's engineers, designers, PMs, whatever, you should do the prototype before, because you're not going to really know how to build the product or what the exact UX should be, particularly like, again, with AI, when these, these new patterns that we can establish without actually having, you know, you play with it, some internal customers play with that, you know, whatever it is. And then when you do the whole formal PRD kind of design process, whatever, You end up writing PRDs and whatnot, uh, and just, and doing proper designs, but because you've done all this pre-work and prototyping, you can actually save a lot of cycles, uh, and just move faster at the end of the day, and then ultimately get a better outcome.
AI assessment note: “Short answer is no. I don't think you lose the design function”
Answered raw tape
D 5 · C 4 · P 4 · Cm 4 4.30
Q What's the biggest product mistake that you've made, and how did you learn from it?
A Yeah, so there's a few things. I'll pick, one of them was the one I kind of just said right now around vault and, you know, doing extraction. Another one that we were doing was, again, at Harvey because it's recent, last year in QFOR, we were revamping our core assistant product. We introduced basically two modes in assistant. One is kind of assist mode, which is doing, you know, Q&A and analysis, and then another one is a draft mode, which helps you draft contracts, draft You know, clauses, emails, you know, whatever lawyers do. And the thinking was you want to give lawyers some choice of like, okay, what am I doing? And pick the right user interface and the model for the job, because those are two different models and user interfaces. Again, this is the joys of user testing. We didn't user test this that much. Um, we should have done it a little more. We were just trying to move fast and get something out as we built it. We launched it, rolled it out and. Almost everyone was super confused on when to use what they didn't know when to use assisted draft. There was, I think it's maybe lawyers, maybe something else. They just didn't know, Hey, when I am, you know, creating a response to a complaint, what should I use draft or assist? And this is like, everyone had this complaint. And then, you know, we started solving this by enablement saying the model, this model was good for …
AI assessment note: “Almost everyone was super confused on when to use what”
Answered raw tape
D 4 · C 5 · P 4 · Cm 4 4.30
Q On the post-mortem side, how quickly do you know when a new product or feature is not working versus it just needs time to sink into customer behaviors and adoption?
A This is something I think a lot about, particularly for our user base, because law firms don't like to move fast or reverse the change. At least the The like admin IT teams are versed to change. And so we can have a product launch and it's not GA to roll out to a hundred percent of users for a few months, um, because people are just slowly adopting it. Um, so I think you, you need to let it bake depending on the customer base. The ideal thing, you know, what we do is you branch out user testing to concentric circles and that your conceptual circles get bigger and bigger. The first thing we do when we build something is. We have an extensive set of lawyers internally doing various roles. We give it to those lawyers and they test it and they give feedback and we iterate. And then we have selected design partners that are consistent for most features externally. We give it to them. They feel part of the process. That's really good. They give feedback. And then there's a broader list of beta testers, like beta, beta customers. We give it to them, iterate, and then there's like the whole, whole general public. And so, you know, this is probably true for most, most products and most enterprise teams is you, you want to make sure you're testing this way, like much more often with, with kind of expanding circles.
AI assessment note: “you need to let it bake depending on the customer base”
Answered raw tape
D 5 · C 4 · P 4 · Cm 4 4.30
Q What's the biggest product mistake that you've made, and how did you learn from it?
A Yeah, so there's a few things. I'll pick, one of them was the one I kind of just said right now around vault and, you know, doing extraction. Another one that we were doing was, again, at Harvey because it's recent, last year in QFOR, we were revamping our core assistant product. We introduced basically two modes in assistant. One is kind of assist mode, which is doing, you know, Q&A and analysis, and then another one is a draft mode, which helps you draft contracts, draft You know, clauses, emails, you know, whatever lawyers do. And the thinking was you want to give lawyers some choice of like, okay, what am I doing? And pick the right user interface and the model for the job, because those are two different models and user interfaces. Again, this is the joys of user testing. We didn't user test this that much. Um, we should have done it a little more. We were just trying to move fast and get something out as we built it. We launched it, rolled it out and. Almost everyone was super confused on when to use what they didn't know when to use assisted draft. There was, I think it's maybe lawyers, maybe something else. They just didn't know, Hey, when I am, you know, creating a response to a complaint, what should I use draft or assist? And this is like, everyone had this complaint. And then, you know, we started solving this by enablement saying the model, this model was good for …
AI assessment note: “Almost everyone was super confused on when to use what they didn't know”
Answered raw tape
D 4 · C 5 · P 4 · Cm 4 4.30
Q Dude, it is so special, Stuart, in person. Now, I wanted to start, when I was chatting to some of our mutual friends, I heard that you early on made a decision to move from engineering to product. I wanted to start with that, actually. Why did you make that decision, and how does that decision lead your advice to others who might be doing the same?
A One big thing I was, you know, always optimizing for is, like, skill set growth. What are the skill sets you need to know? What are the skill sets I'm good at? What are skill sets I want to get better at? Versus, like, is my title a product manager? Is my title a software engineer? So, skill set growth was, you know, super important. I think one thing back then was like, you were more bucketed into different roles than you are. So focus on skill, skill set growth. And, you know, early on, I think I had read Sam Altman's post on how to be successful. This was before Sam Altman was Sam Altman. And one of the things that really stuck with me was working hard and focusing on something that you're good at and like getting better at will compound significantly over time. He said in that, in that post that. You want to try to strive for the top one percent of what you do. And of course, not everyone's going to get that, but if you have that mindset, then you'll get there. And so I looked at myself and I was like, what do I actually, you know, want to do? What am I good at? What am I like intellectually interested in? And sure, software engineering, I had done a CS degree at CMU. I was above average probably at software engineering, but I realized I didn't really want to get better. But instead I wanted to get better at things like, you know, commercial sense, leadership, product sense…
AI assessment note: “I realized I didn't really want to get better. But instead I wanted to”
Answered raw tape
D 4 · C 4 · P 4 · Cm 4 4.00
Q Yeah, I think your willingness to open up is probably so much greater when you're speaking to the An LLM. Uh, final one. Why do agents need humans more than humans need agents?
A There's this one interesting thing about Gen.ai where, so the, who is the consumer of the product or service or the work that's being created and who is the producer of it? You know, examples here are in architecture, it's the client versus the architect creating the design, or there's the marketer versus the brand person for content, or it's like the client versus the lawyer. And do you end up like, what model do you end up taking? Do you help the producer, Tenex their output and make them more productive or go direct to the consumer? And a lot of people have thought they can go directly to the consumer and say, here's a bunch of leads, or here, but here's a bunch of, you know, legal work, whatever. But ultimately, I think the humans don't just always trust AI. They trust other humans using AI. I think, again, depending on the domain, the right thing is more likely that the agent partners with the producer of the work to deliver for the consumer.
AI assessment note: “humans don't just always trust AI. They trust other humans using AI.”
Answered raw tape
D 4 · C 4 · P 4 · Cm 4 4.00
Q Yeah, I think your willingness to open up is probably so much greater when you're speaking to the An LLM. Uh, final one. Why do agents need humans more than humans need agents?
A There's this one interesting thing about Gen.ai where, so the, who is the consumer of the product or service or the work that's being created and who is the producer of it? You know, examples here are in architecture, it's the client versus the architect creating the design, or there's the marketer versus the brand person for content, or it's like the client versus the lawyer. And do you end up like, what model do you end up taking? Do you help the producer, Tenex their output and make them more productive or go direct to the consumer? And a lot of people have thought they can go directly to the consumer and say, here's a bunch of leads, or here, but here's a bunch of, you know, legal work, whatever. But ultimately, I think the humans don't just always trust AI. They trust other humans using AI. I think, again, depending on the domain, the right thing is more likely that the agent partners with the producer of the work to deliver for the consumer.
AI assessment note: “humans don't just always trust AI. They trust other humans using AI.”
Answered raw tape
D 4 · C 4 · P 4 · Cm 4 4.00
Q You're building one of the hottest AI products in Silicon Valley today. I'm fascinated to hear your thoughts on this. Cursor is the golden child. There's also Codium. To what extent is their lock-in for Cursor, and how do the two compare in your mind?
A I think it's like UX is super important. You know, I've heard some engineers say they like Cursor's agent mode much better than, than Winsurf. Some people have said the opposite. So it's like, What is the most delightful UX and short, maybe you can copy UX, but I think it's, it's a little more nuanced and then it's like data and governance, particularly in the enterprise. Like the models are only as good as the context you give it. And so any of these products, can you actually tap into knowledge that the, that the enterprise has about its code, its products, its, you know, architecture and use that in more helpful ways. To, you know, tailor the outputs that do engineers write. And, you know, I know Codium has, has focused a lot on getting a lot of the enterprise data, the enterprise understanding of existing code bases and making sure those, that, that like knowledge architecture is really good.
AI assessment note: “I know Codium has focused a lot on getting a lot of the enterprise data”
Answered raw tape
D 4 · C 4 · P 4 · Cm 3 3.85
Q Dude, it is so special, Stuart, in person. Now, I wanted to start, when I was chatting to some of our mutual friends, I heard that you early on made a decision to move from engineering to product. I wanted to start with that, actually. Why did you make that decision, and how does that decision lead your advice to others who might be doing the same?
A One big thing I was, you know, always optimizing for is, like, skill set growth. What are the skill sets you need to know? What are the skill sets I'm good at? What are skill sets I want to get better at? Versus, like, is my title a product manager? Is my title a software engineer? So, skill set growth was, you know, super important. I think one thing back then was like, you were more bucketed into different roles than you are. So focus on skill, skill set growth. And, you know, early on, I think I had read Sam Altman's post on how to be successful. This was before Sam Altman was Sam Altman. And one of the things that really stuck with me was working hard and focusing on something that you're good at and like getting better at will compound significantly over time. He said in that, in that post that. You want to try to strive for the top one percent of what you do. And of course, not everyone's going to get that, but if you have that mindset, then you'll get there. And so I looked at myself and I was like, what do I actually, you know, want to do? What am I good at? What am I like intellectually interested in? And sure, software engineering, I had done a CS degree at CMU. I was above average probably at software engineering, but I realized I didn't really want to get better. But instead I wanted to get better at things like, you know, commercial sense, leadership, product sense…
AI assessment note: “realized I didn't really want to get better. But instead I wanted to get better”
Answered raw tape
D 4 · C 4 · P 4 · Cm 3 3.85
Q To what extent was the product requirements the same when shifting from ancillary to ancillary? For you leading product, how easy was it to be plastic across multiple different verticals?
A Yeah, so it was hard, I'll say. I think for vision, like, three D and two D, um, it was more straightforward. You know, scale was very operational, and so you'd have to tweak the operations to ensure quality for warehouse robotics versus self-driving. That was, like, Fairly different. But when we started getting to outside of vision, we had to build brand new products. Like I started the e-commerce team at scale where we were labeling e-commerce data, like from, uh, Meta, from Instacart, from DoorDash. That was a completely different product than, uh, the core two D and in three D. And so, yeah, you, you do have to build products, but I think, you know, credit to Alex and his tenacity for being able to say, guys, we gotta do it. It's fine. It's gonna cost a lot of thrash, but there's these huge markets out there. One thing I'll also say just on back to the markets thing is conversely, great markets can hide execution problems. An example with this is Uber. So there was delete Uber. There was a lot of internal chaos at Uber. So I, I interned at Uber in. Like way back when it was super fascinating to see that this was when like, this was kind of like their heyday to TK and everyone was like really going at it. And 2019 or whatever, there's like that deleted Uber. There's a lot of chaos. But Uber is standing, and it is profitable, and you know, obviously credit to Dara, but the ma…
AI assessment note: “outside of vision, we had to build brand new products.”