Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Yeah, and Kevin, to build on, on Ryan's point, I would love to hear what your own process was to figure out if this new set of enabling technologies were right for Sprig, and at, and at what point in the early days of open AI's models or, or others?
A When we were trying to figure out how to use it, uh, we basically did the same thing we were doing before when evaluating other open source models. You know, had a, a set of what we call Source of truth data. Um, you know, surveys from customers, customer feedback that we'd have to feed through it and evaluate and see, you know, how, how good the model was at, at performing its, uh, its function there. Um, what really kicked off our decision to, uh, to promote this model and to use it in production was, you know, a discontinuous jump in efficacy that we saw. So it was significantly better than what we were using before. And we determined that We'd be able to use it at that high level of efficacy ongoing in the future, you know, assuming we had some ways of being able to continuously test it and make sure it was still able to deliver over time.
AI assessment note: “had a, a set of what we call Source of truth data”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Switching gears slightly. You both mentioned this a little bit, but can you share more about the org structure has changed and maybe product teams have changed if it's been significant over the past year?
A Um, I can start with, uh, specifically. AI and how we've, how we developed it at the company. I think recently the past six months we've modeled our, what we call our AI squad based on some successful examples we've seen in industry, kind of like the vertically integrated approach. I think Ryan mentioned this, but a dedicated designer, dedicated product manager, dedicated engineers that are responsible for full end to end delivery of, you know, AI related features. I think this is in contrast to some other examples. And I think something we've tried before where You know, maybe you have a dedicated ML team or AI team, ML scientists, ML engineers, research engineers who will, you know, sort of receive tasks and return the results. You know, maybe you have a separate product team that needs AI, uh, some, some AI feature in what they're building. Uh, they'll kind of contract it out in a sense to this separate team. You know, I, I've read a lot about companies that have kind of taken that approach and from the times we've tried it, not as much success. Uh, I think mostly because there's an alignment issue sometimes, um, it's just kind of harder to get that end to end release of a quality feature. And so, like I mentioned, yeah, we've had a lot of success with this fully vertically integrated approach.
AI assessment note: “we've modeled our, what we call our AI squad... vertically integrated approach”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q How did you, or how do you set up this sort of testing environment or approach where you can look at a different model and decide if it could be useful for Sprig or not?
A Yeah, um, now it's a bit different. Um, you know, I think prior to large LLMs becoming available, you know, our testing framework was asking a model, maybe a series of specific questions like, hey, does this response represent somebody talking about pricing? Yes or no? Is this person happy or sad? Things like that. You know, with the advent of chat completions, um, and how, you know, even classification tasks are going through chat completions now, um, we've had to sort of shift that to Giving an input into an LLM, gathering the outputs, and then either semi-automatically or manually determining if the output is what we're looking for or not. Uh, one example of, of, you know, something we've been doing is feeding in data from different types of questions being asked. You know, maybe question one is, how's your experience? One to five. Question two is open text. What else can we do to improve your experience? Um, and we've been experimenting with feeding or have been successful with feeding, uh, results from both of those questions into a model. And then asking it to, for example, create correlations between those questions. For example, of the people that gave it a four or five star, what are some of the interesting themes that we see in their open text response? And that's just, that's something that's really hard to evaluate in an automated manner, at least right now, because…
AI assessment note: “shift that to Giving an input into an LLM, gathering the outputs”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Kind of maybe continuing to go down a little bit of this path of, of product development and prioritization kind of in this world of large language models. How has the hype and euphoria in the market changed the way that you think about any of these prioritization decisions?
A Yeah, I think we've noticed a couple of things. Um, like you mentioned, there's the obvious, uh, sort of external pressure. I think with the advent of some of these models being released, the bar has been lowered. Quite a bit. Although some of these analyses or things you want to do are, you know, more achievable. And so at least I think from our point of view, that's kind of prompted us to honestly, to, to work faster and to build, build better and faster. On the flip side, I think it's allowed us to build more than before. Um, you know, and as one quick example, I spent, you know, several years tuning that, that very first sort of feature we built, um, a lot of effort spent on that, not much effort spent elsewhere because we really wanted to focus and make sure that was You know, where we were spending our effort, but now that the floodgates are open, we can work on a lot more stuff that expands our ability to, to, to produce.
AI assessment note: “prompted us to honestly, to, to work faster and to build, build better”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q It's a really interesting way to frame it. So to pull on that a little bit, how did you answer that or sort of figure that out in the first few years of Sprig?
A Let's see. Well, maybe you can back up a little bit, um, and say that's actually not where we landed to begin with. Um, you know, I think at the very beginning, we tried to follow the mantra of, you know, keep it simple, try the simple thing first. If that doesn't work, then you increase in complexity. You know, I think the very first attempt we made at this, uh, which we thought could work because at the time we had a relatively narrow customer base and a narrow data set, uh, narrow data domain. Um, and that was basically like a, a categorization, uh, sort of approach. So, you know, of, of all the things you could ask somebody about their use in a product, maybe there are 10 big primary themes you could talk about, you know, ease of use, user interface, pricing, support, whatever. And maybe within each of those, you have some subcategories, so that kind of a hierarchical categorization. And we found pretty quickly, even for our narrow subset of, uh, initial customers, that this just didn't capture everything. Like I mentioned, language is very nuanced, but also people's needs and thoughts about how they're using a product are very nuanced. You know, somebody could be talking about the, you know, maybe they want cheaper pricing for the product and somebody else could be saying, you know, they want team pricing instead of individual pricing. Uh, a really basic analysis would put…
AI assessment note: “the very first attempt we made at this... was basically like a, a categorization”
Answered raw tape
D 5 · C 4 · P 4 · Cm 4 4.30
Q Do you think it, it means that these large language models are, quote, less intelligent than a lot of people think?
A I think they're less capable, maybe is a better word. Yeah, it's just, uh, They're built, the way they're trained supports a handful of really important use cases, but just not everything out there. People have found ways of getting around this problem using LLMs. Like, for example, asking your LLM to write Python code that will add to 16 plus 32. Um, and in that case, it's really good at it. Uh, but, you know, you don't necessarily want to have to go through that circuitous route to get your answer. Aside from that, you know, I think We're, you know, related to that, we're starting to see the down part of the hype cycle with this, uh, just the limitations that become, became pretty obvious. Maybe that that quantitative sort of failure falls under a broader category of failures that these LLMs will experience. But, you know, despite not knowing the answer, we'll give an answer very confidently, right? I think that's gonna be a major thing to over, a major challenge to overcome because You know, right now there's really not a good way to evaluate these at scale or auto in an automated way. You can't necessarily read through every single sentence a LLM generates when it's summarizing text for you. And I think that's going to hurt trust, uh, in a lot of these and maybe kind of slow down adoption, because if you're really counting on these, these models to, you know, sort of replac…
AI assessment note: “I think they're less capable, maybe is a better word.”