Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q just doesn't work, and so, like, sometimes I'll be using my phone, and I'll use the IOS keyboard transcription to type in the field and then, like, say a bunch of stuff and then send it off. But this suggests to me that consumers really want voice mode that works, and yet it's just not working yet for the major LLM apps or for anyone. Why is it working yet?
A It is pretty hard to do, because you want, you want two things. You want, uh, you want to be able to say things that you want, but you want sometimes for it to execute it, sometimes to wait for you to, like, finish and add something in the sentence. Sometimes you want it to be interactive, so it asks you questions back to clarify and get some of the additional detail, and all of that is actually pretty hard. Like, that's where kind of the, the magical, like, ideal version of a voice agent for us comes through, where you need the speech-to-text element, you need the transcription side, unique You need then the kind of the turn taking mechanism. So like, when do you finish sentence? When, when is it likely based on silence, based ways likely on the context? And then sometimes you want it to speak back and clarify, or at least give you the text back to clarify, and then maybe execute set of instructions. So that problem is still very hard research. So I agree with the claim that like this orchestration side has not like passed a true conversational agent Turing test. Where it like behaves as you would expect from another person where you can say.
AI assessment note: “It is pretty hard to do, because you want, you want two things.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q So when you say can bring convert conversational agents is the biggest priority. Is this for customer service type use cases? Like what are the most popular use cases for conversational agents?
A Yeah, like we want to be a partner for like full interactions between business and businesses and their customers or their audience. Um, I'm saying that the audience because that will apply in support. Support is the easiest one because that's where it's most ready, but like, and that's maybe the big difference to how we see ourselves to some of the other companies in the space is this can also apply to sales. You can, you can have the proactive side of reaching back. You can have AISDR versions of that. Yes. Um, and then you can have all the way to the marketing use cases, where we are your partner for, for working on, on, on, on, even outside of like the, the conversational agent space of how you create a great marketing campaign.
AI assessment note: “Support is the easiest one because that's where it's most ready... also apply to sales”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q It feels like a big part of the magic of 11 was your voices were much more human sounding. How did you accomplish that?
A Kind of give you a, um, a quick, quick synopsis of what, how we think about the models on the text to speech side today. In any model you need, you need the architecture, you need compute, you need data. So architecture innovations were one thing. The data part was the second big thing. With audio, you will have, um, you will have a lot of audio data available, but frequently you will not have it annotated in the right way. You won't have which speaker is speaking when, um, some of the what Uh, is annotated, but the how isn't. So, like, as we are speaking now, what's the emotions that we use? What are the actions that we use? So we would invest a lot internally on effectively creating our own data labelers, our own team, to be able to create those data sets that will be better. And that was a combination of, of course, like, semi-automatic techniques, and then, and then, and the manual techniques. And actually, a lot of the models that we did afterwards actually spun out from a lot of that research, too. So speech-to-text model, Initially it was a model we did for ourselves because the models on the market just weren't good to annotate that data. And then another brilliant researcher on our team was kind of being able to construct it so we could span it out as a model that we brought to the customers.
AI assessment note: “we would invest a lot internally on effectively creating our own data labelers”
Answered raw tape
D 5 · C 4 · P 4 · Cm 4 4.30
Q It was like awful dubbing, you know, happening in Poland previously. So that's like one example of the, you know, the second order effects. What are the other second order effects you're seeing of ubiquitous good text to speech, speech to text? It seems like across a broad array of languages, because whatever about in English, just this didn't exist in Polish or Irish or, you know, pick your language.
A One, like, breaking down the language barrier, we, you know, the, kind of, the inspiration came from, from the movie side, but it also applies in any, in any communication setup, like, could in the future, could I travel to another country and speak, speak Polish or speak English, and that, that language isn't being understood in the local native language. Like, from Hitchhiker's Guide to Galaxy, this version of the Babelfast, exactly, that you can, like, actually understand the world. And voice, of course, will be an interaction layer, but similarly, all of us will have our own, Kind of extension and voice agents that can help on, on our behalf. And there is like very clear and, and, and, and great examples of that of people that lost their voice and can get it for the first time, for the first time back. We see that everywhere, whether that's people that lost it due to ALS or throat cancer that can get it back. Uh, just recently there was an example of a patient that had Neuralink and worked with them to bring the voice that that person could speak with their own voice back to the, Back with the, with the family around, we worked with, with, um, with, with the lady that lost her voice before, before she got married, and, and then finally technology became possible. We, we were able to recreate that voice, and for the first time she could replicate the, the marriage ceremony a…
AI assessment note: “One, like, breaking down the language barrier... people that lost their voice and can get it”
Answered raw tape
D 4 · C 4 · P 4 · Cm 4 4.00
Q in talent or something like that, is it that you are building your own software where other companies might have bought software like a Workday or a Greenhouse or something? Is it that they Are using the existing software you have better? Is the process that will be spreadsheets in a traditional company are built with software? How do you kind of use the software in these sorts of organizations?
A Yeah, like, there are sometimes, but we still use a lot of, like, the, the traditional vendors. Like, one pattern is, of course, elementifying everything, like, making the, the data explorable for you to be able to interact with it, like, who's in the pipeline, what worked, who does the best references, like, all of that, all of that works, so you can double down on that. But two, it's frequently things that you manually do, that a lot of the current, like, there is a gap between where, where, where the agents are today versus what you could do if you have the technical skill set. And a good example is, like, How do you scrape all the right profiles to be able to reach out to the right candidates? Um, so you, like, analyze whether it's, you know, how much I should want to say, but, but, uh, the, like, try to detect specific things that we know worked, so you bring that across to the, to the, to the, to the people. On go-to-market side, like, there's just so many things you can do with, with, with additional amplifiers. You know, it goes from Understanding what case studies are relevant and creating a good pre-read for you before you go to the meeting, through creating the AISDR experience that we spoke about, to creating an entire deck experience, so you have like a pre-populated deck with the right numbers that is customized to that customer, which you want still the person to…
AI assessment note: “we still use a lot of, like, the, the traditional vendors.”
Redirected raw tape
D 2 · C 5 · P 4 · Cm 4 3.70
Q Okay, but are there interesting differences beyond, like, correlates like size?
A What I can say is, like, slightly different to your question. The people interacting through voice and the performance we see for, like, how they interact with, with, with, with the business changes just by nature of interacting with voice. A good example, you can contact 11 Labs and register for For your interest, you go through the form, and at the end of that, ah, we have supplemented that, that instead of going through the form process, you can speak with our agent and leave more details. And what happened are two things. One, people were actually much more keen to leave the forms through speaking with the agent, so we would go through the form a lot, a lot easier. But second, they would be a lot more open-ended in terms of what the use case are. So they would start giving us information about the wider set of use cases, the complexity of the use case. So like the writing out was tedious and, and, and, and tricky.
AI assessment note: “What I can say is, like, slightly different to your question.”