The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Des Traynor argument clarity score 4.4/5 from 16 exchanges on raw tape · average scores: directness 4.9 · coherence 4.6 · precision 4.2 · compression 4.1 record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
16exchanges match
16on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q of UX? Um, so you want to give the, the, the users, or your customers, users, um, you know, a choice in terms of like how they navigate and how they get to the answer. Um, is the idea that, uh, you can almost like pick AI versus a human. You can fast forward to human directly. Uh, how does that, how does that all work and any early lessons?

A In our manifesto, we say, like, we believe the future of support will be humans plus AI. And what we mean there is the AI will, uh, so basically some conversations will be answered entirely comprehensively by AI. How do I reset my password? Here's how. Uh, how do I get a refund to click this link? That type of thing. Just a complete answer. Customer gets exactly what they want and they leave. The next set will be things that the AI will attempt to answer, but ultimately might fail over to a human. And then some, the AI will just be like, I'm not touching that. It sounds like a sales query. I'm going to hand that straight over to a human, right? So that's the first piece. The second is there's a symbiotic relationship between the, the, like the bots and the humans, if you like, right? The, or the AI and the human, which is when are, you know, humans are in the inbox, the AI can help them. It can do things like summarization. It can like, you like, look stuff up for them and all that sort of stuff where we're investing a lot there. But then also the humans help the AI because when, when you, you know, one of the features of Finn is you can go through all the answers it's given and be like, oh, you got that one wrong. Let me teach you. And you can give Finn new facts to learn from so that it doesn't make its mistakes again. So I think you'll see this kind of relationship where lik…

AI assessment note: “we believe the future of support will be humans plus AI”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q And by the way, uh, what is the latest number in terms of resolution rate? I think I've read somewhere. 41%, uh, of queries. Is that higher now?

A It's a bit higher. It gets higher every, every few weeks. We, so we just shipped Finn earlier this week in 40 more languages and that will kind of give it a higher involvement rate and probably a higher success rate too, because it was probably, you know, speaking bad, bad language, like, you know, bad English or whatever in previous cases or whatever. Uh, so that number is creeping up. I don't have the current one, but like forty-ish is, is roughly where we're at. I will say like that, That kind of hides, like on average, I would be confident any business who turns on fin will get at least like 25, 30. We have a lot of people getting 70 and 80. Uh, it kind of depends on the simplicity of the support function that you're staffing. So if like you have some sort of Pareto style, 80, 20 for your support queries, as an example, like utility provider deals with like open an account, close an account, change an address, register a meter reading, whatever, uh, those four or five questions often account for 80, 90% of the entire inbound. In those cases, FIN delivers exceptional results, as you guessed, because you can target, you can deliver so much value by just getting really good at four or five things. Um, so sort of like it's, the 40 thing is like, is definitely like, you know, it's very, very real. There are just, there's a lot of cases where if you have a simple, simple enough s…

AI assessment note: “It's a bit higher. It gets higher every, every few weeks.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q And for Finn, you charge, uh, based on resolutions. What is that concept? Uh, and is that something that customers understand or need to be educated about in this brave new world?

A Um, there's a tiny bit of education. Like we actually consider resolution the same way they do. Uh, but you just have to explain a little bit. So, so let's say a thousand conversations come into a business. Uh, let's say Finn only touches 700 of them because it looks at 300 and goes, I don't know what to do with that. Uh, because maybe it hasn't read the right docs or hasn't been fed the right information, or maybe they're gobbledygook or spam or whatever. It doesn't really matter. It's 300 of those. It's not touching. That's not relevant. So the first thing we would quote is our involvement rate, which would be 70% in this case. 70% of the conversations Finn jumps into. That's not what we prize for. We prize for when Finn has given an answer and the customer has either explicitly said they, has either just closed the messenger and gone on and done the thing they wanted to do, or has explicitly said that answered my question. The only thing we don't charge, the only time we won't charge here is if the customer pushes back and says, that's not right. This is wrong. And we hand over We're human. That's what we don't charge. Uh, we basically charge when we gave the customer an answer that they saw and they didn't have any follow-up questions, which is exactly how CS reps are measured as well. And that like, no one goes chasing people being like, are you sure? Are you sure? Are you…

AI assessment note: “Um, there's a tiny bit of education. Like we actually consider resolution the same way”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Yeah. And what makes it so that it's in the kill zone?

A Good question. It's basically that large language models are really good at conversational back and forth, uh, What I would call a give this, answer that type workflow, which is read this and answer the following question. Um, and disambiguation, just generally speaking, what do you mean? Clarification, the pushing and prodding. A lot of previous models we played, it would give up or tap out somewhere and it'd not be useful. Whereas ChatGPT was both pretty quick and very good at like conversational problem solving effectively. That's what a lot of like, not all customer service, but there's a good chunk of undifferentiated customer service that falls into that area. Stuff like, how do I reset my password? It's not like a brand building opportunity from which you can establish a long-term relationship. It's just a very answer the damn question. And the best version of that question is not an artisanal hand typed, well-crafted, beautifully worded, eloquent letter. It's here's the link to reset your password. And people actually care about instancy more than they care about like, uh, you know, let's just say the personal tone in those types of things. So, uh, so I just think Does it, it was, it was very obvious at the very start, there's a certain chunk of the work that's definitely doable. And over the last while, we've just been expanding that chunk to include things like, you k…

AI assessment note: “large language models are really good at conversational back and forth”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Okay. Very cool. What does voice fit in? I saw you made announcements around Intercom phone. Is that, is that, is that in like a voice version of Finn? Is that different?

A No, actually, well, the first thing we have to build, because like, there's two things everyone needs to know about Intercom. One is that we are a complete customer support platform, and then secondly, we have the best AI. And we ran a funny billboard about that, but the reason we had to build phone was just because people wanted to, people wanted to consolidate onto Intercom, but they couldn't because we didn't have a phone solution. So the first thing we do is just build a phone solution. But the question you're talking at, which is the one everyone's excited about, is like, will Finn answer your phone calls? And the answer is like, not today, but absolutely voice is like, voice is still a huge amount of support volume. Phone calls or like, uh, or like increasingly with like, uh, Gen Z voice notes as well are becoming quite popular as well. Um, we've played with everything you'd guess, Synthesia and their synthetic bots. We've played with 11 Labs and their synthetic voices, et cetera. And in general, there's a, there's an obvious path that we're going to get to here. I'd say the, the reasons we're not building it right now, uh, one is, is like, you know, I think it's not, it's basically not customer demand number one, but it's, it's, it will get there for sure. Uh, It's a, the latency is a little bit of a problem in a phone call, and like as in a good GPT-IV query might take …

AI assessment note: “will Finn answer your phone calls? And the answer is like, not today”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Another really interesting aspect of the journey is that, um, you guys have been bi-continental and distributed from the beginning between, again, San Francisco and, and, and Ireland. Um, where, where do you stand today on the whole distributed versus centralized work from home versus work from the office? Any, uh, anything that you've experienced that you could share?

A So you're correct. We've been distributed forever. So like when, when the world went remote in 20, uh, March, 2020, was it? Uh, we weren't that shocked and not like me and Owen at the very least. I always had a remote relationship. He was always in San Francisco. I was always in Dublin. So it wasn't like a big, oh my God, what do we do now type thing? But I, I, I think, um, you know, so like in some sense, distribution or distributed offices was always fine for us. Uh, we had those muscles pretty warm. The going remote, I did not like personally. And then separately, I think, uh, we were now like two days a week in the office. And I think, um, When I look around the days in the office, there's a lot of like, you know, positives to see. There's people smiling, high-fiving, you know, there's people going for drinks after work. There's people having like standup meetings by a whiteboard, really chasing down like tech specs for what we're going to build. And all I see in all that is the stuff that wasn't happening when no one went to the office. So I think like, I'm a big believer that like face to face matters, not just for collaboration. It does matter for collaboration, but also I think it matters a lot for, uh, Just for like, you know, work is such an important chunk of everyone's lives, whether they want it to be or not. The idea that you could do it and have like basically no…

AI assessment note: “we were now like two days a week in the office.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q the Intercom website at intercom.io slash believe. You basically have a manifesto which uh, seems like it was signed by just about everyone in the company if I, if I, you know, from, from the looks of it. Uh, so was that that it was like, like this big rallying cry where, where, um, you know, uh, you all decided that the whole company was evolving towards this new goal?

A That's exactly what it was. It's like the manifesto is literally what we believe about the space. And it took us quite a while. Like it took us quite a while to distill these things because so much was changing. Uh, but you know, we, we have our sort of what we call our big beliefs, which is that like CS is in need of a significant upgrade. Um, That like, you know, uh, AI is going to be the thing to do it. There'll be winners and losers, et cetera. And, uh, like we've been distilling this manifesto for like two years as we've been kind of progressively learning what AI will and won't do and where humans are involved and aren't involved. And, um, and ultimately us putting it out there was us where our way of telling the industry, right, if you choose to go with intercom now, this is what, this is what underpins everything we build, everything we do is set to advance this idea here. So it was really a good kind of decrystallization after maybe a year or a year and a half of consideration of what the hell we're all about.

AI assessment note: “That's exactly what it was. It's like the manifesto is literally what we believe”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q And, uh, AI first, um, Is a question of like how much development effort you spend on AI or is basically what you're saying that you're completely flipping the logic of customer service, meaning that people's interaction, initial interaction will be with AI first and then humans will be a fallback?

A We have this idea that like, uh, if AI can answer the question, uh, it should. And we spend a lot of time making sure that AI, our AI finn is aware of what it can and can't answer, how certain it is. The logic for that is quite simply an instant answer is almost always what the customer wants and they will sacrifice a lot for an instant answer. And if you look at your own online search behavior, you go to a website, if you don't get an answer to the thing you're wondering in like four seconds, you close the tab and go back to your Google search and jump down a link or whatever. Like that's just the way customers behave. So instancy is a really valuable thing. Uh, so we, we say AI first from that perspective, but there are a few other perspectives that we think about AI first from like when we're building something or we're trying to capture a new workflow, let's say we're rebuilding our reporting. The mindset now is, can AI do this? So let's say a classic customer support leader job might be like, what new issues are occurring in our, in our, in our volume of like support at the moment? There are loads of ways to do that, but right now we think about how can we do that from an AI first perspective? How can we get out, get ahead? How can we preempt the desire and actually automatically surface the answer before it's even been asked? We, it's very much both, uh, how we think abou…

AI assessment note: “It's very much both, uh, how we think about our support model”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q And, uh, just, um, going back to the, the, the, the question of the fundamental, uh, foundational model that you use. So where are you in that exploration of, um, GPT-IV versus Claude versus Llama versus building your own thing, and how do you think about it?

A Mistral, blah, blah, blah. Yeah. I mean, I, I think, um, Look, we care a lot about using the best product out there. We don't, we don't feel the need to be like checking it every week because, you know, frankly, you might forget to build the rest of the software that I, uh, enumerated earlier. But, um, we have a torture test of like thousands of questions and scenarios that we run any given model through. To run it through is, you can do it like in a reasonably automated way, but there's a lot of like human eyeballing to really make sure that we're We're assessing things correctly too. Uh, so given a model, let's just say Anthropics, Claude, and Mistral, and Llama, and OpenAI-Four, and whatever we might roll ourselves, we look at their performance against all of these things. And we, it's a, there's obviously like trust, reliability, accuracy, all that sort of stuff are like, they're our first concerns. Can it actually do the job? The secondary concerns then become things like speed and like cost and all that sort of stuff as well. At the moment, like we're on four and we're happy there, but we're not, We have not moved into an optimization mode here yet. Like as in at some stage, somebody might talk to us about anything from like margins to speed or performance, whatever. And that's where we might start saying, Hey, right. Okay. Given that we've kind of built all the stuff we …

AI assessment note: “At the moment, like we're on four and we're happy there”

Answered raw tape D 4 · C 5 · P 5 · Cm 4 4.55

Q What are you learning so far in terms of like how people react to AI, right? So seeing from the perspective of users, you know, I think there's a little bit of a PTSD that accumulated over the years with, you know, chatbots. That never gave you exactly the right answer. Um, so how you, how do you overcome that?

A Yeah, you're totally correct. So like, there's like been generations of shit chatbots basically. The first generation was like, was, uh, was like just button powered bots. Like, are you trying to do X or Y? X. How are you trying to do X? One, two, three, one. Like, it was very much like a phone tree, but in a messenger. Um, that was gen one. Gen two was then a little bit of AI where it's like, type your query and I'll try and guess what you're, what you're trying, what you're saying. And that was the gen two, and that's what our resolution bot product was. And in this new generation, I think people's behavior has changed a little bit for two reasons. One, we're like, you know, 10 years into it, but two, like ChatGPT and then obviously Gemini, and who knows what will happen when like Gemini goes on Android phones and Apple finally launched their thing. But I do think people are starting to realize bots are good. Like, as in people play with ChatGPT and their expectations are growing. So when we first launched and people thought we were a bot, they, you know, people drop immediately into what we call bot speak. So it might be like, hi, Matt, I, I, I ordered the t-shirt two weeks ago, and I'd like to, and as soon as they see a bot replying, they say t-shirt refund, please, you know, because they've given up on all the English or whatever. They talk machine to it. They talk machine…

AI assessment note: “I do think people are starting to realize bots are good.”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q enterprise use cases of generative AI is hallucination, right? It's one thing if, uh, you know, you ask judge EPT to write poetry and they hallucinate. It's a whole different thing. If you have a customer experiencing a mission critical need, um, what kind of work have you all done So far, you mentioned the term safeguard. You mentioned RAG. What have you done so far to limit that problem?

A We, as I said, we have like genuinely a very exhaustive, uh, torture test that we put any given model through that is almost designed to trick it into hallucinating and see where, where it'll fail. We've also done a lot of work in, um, in, in getting the, the models to like report out on their own, um, how would you say certainty of the answer? Um, and we can then, that gives us ourselves a threshold where we say, Hey, like we could have higher coverage, but we would do it at a cost of, uh, more, More like bad answers basically. So we, you know, the things we care about are, uh, first of all, we, you know, we wouldn't go through the same reason we won't go with three, five. We, we take the obligations of our customers reputations quite seriously. So we don't want to have hallucinations. We put heaps of work into that. A lot of the times, uh, when, when, when Finn gives bad answers, it's not hallucination. It's actually usually just bad content and it's stale content on our customers help centers, uh, That's happened or somebody before, like somebody yesterday answered this question and they gave the wrong answer and now Finn thinks it's correct. So it's gone with it wherever. So that's where we have that idea of human helping the AI. So human logs in and fixes it. But there's a lot of work done to sort of guard rail to certainty checks. So like to just threshold how confident F…

AI assessment note: “there's a lot of work done to sort of guard rail to certainty checks”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q obviously is, uh, the holy grail and super interesting. Uh, so without getting into sort of personalized AI, as of now, do you enable personalization? Meaning, uh, if I show back up, uh, three weeks later, you will remember who I am, my history, that kind of stuff. Do you have, like, some kind of, like, underlying customer data platform where you store all of this? How does that work?

A Yeah, great question. So, uh, so Intercom is an actual CDP. We're kind of one of the first, we were out before segments and all those with our version of this. Um, so all our users store custom data on their users, uh, and it would be things like what price plan is Matt paying and how many teammates has he, and how many files has he uploaded or whatever makes sense for your business. All that information is also fed to Finn as well, and Finn can use all that as well. There's obvious sort of, uh, value we get out of that. Like, so if it, if Finn knows that you're a premium customer, it can give you the answer to your question, assuming you're a premium customer. Like, so You know, Spotify has a different set of features for premium users than it does for regular. And if it knows Matt is like a premium user, it'll give Matt the right answer. Um, more generally, like, uh, we think about personalization. I just, I, I'd sort of say there's like, there's levels of automation that we sort of work through. Like, you know, if you can think about this, like the same way we went through, say, self-driving cars, right? You know, in the eighties, there was no automation today. There's a lot of automation where like maybe L four or something like that in support that you go from like no automation to maybe like old school bots to like, um, LLM bots. Dynamic answers will, uh, you know, the id…

AI assessment note: “all our users store custom data on their users... fed to Finn as well”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q Transitioning a little bit into the business aspect of all of this, one question that's top of mind for many people is cost, and then the impact on gross margins if you're a business, if you're a business using Uh, generative AI. What have you learned so far?

A It's definitely a valid concern. Like, uh, an interesting mindset that we've had to adopt in Intercom is like, there are cool features. This is the first time in my career this has been true. There are cool features that we can't afford to build, right? So as an example, like summarize every conversation Intercom powers every month. Uh, that would be summarizing five hundred million conversations a month. That would bankrupt the company in some amount of years, right? Like it's, it's, it's, and we wouldn't be comfortable passing that cost directly onto our users. So it's definitely a, Like how much will this cost is a new sort of a step in our, should we build it? That just wasn't there before. Related to that, by the way, so is latency because like these things are just slower than normal software. So you have to think about where and when you inject delays too. Uh, but so what have we learned aside from that? Like, I think there's a few different things. Um, It's awesome that there are open source models coming up, uh, that we believe we can hand over more and more of the workload to. We haven't really gone deep down that journey yet. Cause as I said, we're not really in a cost optimization mode, but we're definitely growing in confidence that we would, we do not have to pay the, the, you know, the full highest price, uh, for every single call we make that we have obvious pat…

AI assessment note: “there are cool features that we can't afford to build”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q of your path towards a hundred percent, which I assume is the ultimate goal, there's a really important concept of, of actions. The, the bot actually taking action. So that's one thing to send a link to reset a password. It's another thing to basically do it for you. Is that, is that, In the spectrum between science fiction and happening today, where does action, uh, generated by AI stand?

A We have, like, internally, we have demoed this, uh, these capabilities, so it's, it's absolutely going to happen, high certainty, and we will definitely do it. If you recall, like, the sort of, the L-one to L-five of support automation, I think actions is probably, like, one level above what we have for sale today. Uh, we are, we are, we are firmly gonna, like, make that happen too. It, it's specifically important, uh, in certain businesses where, um, You know, like let's just say approval or refund requests or whatever. Like we can't go and do them, uh, without like having a proper, like, um, way to go and ping APIs. Like the, the thing that will for come before actions will be like third party data lookups. So pulling in for information and relaying it back to customers. So like, why was I charged 59 dollars on my credit card? Here's the answer. That type of thing. That's, that's, that's, that's just a read. It's not an action. Uh, but actions are definitely gonna happen in terms of like, you know, it's, it's not that hard. Like today, Intercom has a model as a feature called custom actions. Where you can actually offer, like click this button to do something, which could be like refund my order or whatever. It's the thing we have to unlock. And it's just, again, it's back to this idea of trust, reliability, and guardrails is when are we comfortable letting Finn do that off i…

AI assessment note: “we have demoed this, uh, these capabilities, so it's, it's absolutely going to happen”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q tech ecosystem in general and, you know, B to B enterprise software startups. Everybody's had to adapt. What did you all do? Uh, it seems like you've had, um, you know, a bunch of like senior changes. Owen came back to be CEO. You have a new president. So what did you, how did you all adapt and become fit, which seems to be the, uh, expression of the day?

A I mean, first of all, I'd say, uh, Owen returning was like, uh, Was the catalyst for all of this. Uh, I think like right-sizing costs. So we obviously had to do a riff like every other startup that was out there. Uh, like honestly, like, uh, refining the strategy, creating a lot of focus and a lot of urgency behind one thing. So moving away from being a broad tool to being a very focused CS platform that's, that defines the industry. And then, and then obviously pouncing on AI as part of that. I think we also had like a lot of work to do around changing our go to market motions, making sure that we're actually accessible and affordable and adoptable by companies big and small. Uh, we did a lot of work there as well, uh, changed how we kind of, you know, represent ourselves to the market. Um, spoke to a lot more clear value. Uh, and I think, um, in general, like, The framing I would use with a lot of the startups I've invested in or spoken to is like, you have to be everything for somebody and not something for everybody. And I think that's for us, like a lot of it, we, we want to be businesses first and only customer support platform or AI first customer support platform. I think a lot of that was the strategy. And a lot of that is like now the positioning and the marketing, and even the fact that like we're having this conversation is a byproduct of the focus that I'm created …

AI assessment note: “right-sizing costs. So we obviously had to do a riff like every other startup”

Answered raw tape D 5 · C 4 · P 3 · Cm 4 4.05

Q Yeah. Is a version of that future Rather than the hard swap, some concept of router where in real time you would send different queries to different models?

A It's possible or like, uh, you know, we, I wouldn't rule that out. I also wouldn't rule out like, uh, so like different features go through different, you know, depending on what we're asking, we might have, we might look for the cheaper calls that are more expensive ones or slower ones or whatever. There's also other stuff like, uh, would we let customers choose? Um, would we let customers bring their own? Uh, like the, you know, there's loads of versions of that. Like we've, we've built it in a sort of like a high cohesion, low coupling way where you You can hot swap if you want. This, the, the challenge we have is just, um, we put a lot of work into making sure that we're certain on things like trust, reliability, et cetera. Uh, if customers are going to, uh, you know, hot swap in their version or like the reason we wouldn't just overnight flip from one to the other is because, you know, there's like reputational considerations on behalf of our customers. We want to make sure that, you know, that these things don't ever do anything we would, we don't want them to do. So we put a lot of work into guard railing them and we just, That's why we wouldn't necessarily be running with a bank of 10 of them, because to own any given one of them in any period, it doesn't ongoing work to it. It's not for free.

AI assessment note: “It's possible or like, uh, you know, we, I wouldn't rule that out.”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.