Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 4 · P 4 · Cm 4 4.30
Q Are there specific dimensions that you look at from an eval perspective that are, that you think are most important in terms of how you think about?
A Yeah, evals for, you know, generation are generally challenging because, um, they're, uh, Qualitative and based on sort of, um, you know, the general perception of someone who looks at something and says, this is more interesting than this. And so there is some dimension to that, but I think for speech, like, um, you know, emotion is something that matters a lot because you want to be able to kind of control, you know, the way in which things are said. And I think the other piece that's really interesting is the, um, how speech is used to embody kind of the roles people play in society. So like different people speak in different ways because they have, you know, different jobs or, Work in different, ah, you know, areas, or live in different parts of the world, and that's sort of the nuance that I don't think any models really capture well, which is like, you know, if you're a nurse, you need to talk in a different way than if you're a lawyer, or if you're a judge, or if you're a venture capitalist, you know, very different forms of speech.
AI assessment note: “I think for speech, like, um, you know, emotion is something that matters a lot”
Answered raw tape
D 5 · C 4 · P 4 · Cm 4 4.30
Q Given that like versatility of model architecture, generality of the building block, like what's, what do you do next for Cartesia? You focus on like the headroom, uh, for Sonic and audio, you work on other modalities.
A You know, I'll take that one. You know, we're obviously really excited about the Sonic work because I think, um, it kind of shows the first, um, example of something that we're excited about, which is it's a real time model. You can run it really, really fast. At low latencies, and it's capturing this idea that you want to generate a signal of some kind. So we're going to continue to obviously improve that piece. Also, you know, just generally, um, things that folks want out of speech systems that, um, need to get built that are orthogonal to, uh, the technology piece, which is, you know, uh, being able to support lots of languages and just generally providing more controls. Orthogonal access that's really important for generative models, which is You know, how do you kind of add more controllability in general to the system so you can kind of get the desired output that you want? So that's obviously one focus for us is, is how to kind of, uh, put that piece in. A few things that we're doing that I think in the short term are really interesting. One is, uh, bringing Sonic more on device. You know, you can run the model real time in the cloud. Wouldn't it be cool if you could run it on your MacBook and it ran real time and it, you know, was, uh, just, just as good. I actually have a demo I can show there, uh, that I think is super cool. Over time, what we want to do is what Albe…
AI assessment note: “So we're going to continue to obviously improve that piece.”
Answered raw tape
D 5 · C 4 · P 3 · Cm 3 3.90
Q Yeah. And Cartesia recently launched, um, its initial sort of text-to-speech product, and it's, it's really impressive in terms of performance and how fast you've gotten to ship something really is that performing. Can you tell us a little bit more about that launch and that product?
A Yeah, I think it was sort of a natural, um, you know, transition for us To kind of now start thinking about how to put the technology to work because, you know, there was a lot of pre-work that happened and Albert continues to do the pre-work for the next set of things. But, uh, um, I think it's sort of like, how do you kind of build an efficient system that will allow you to do, for example, in this case, voice and audio generation. So I think the, the way we're thinking about it is we're building these fairly general models inside the company that allow us to kind of do, uh, fairly generic tasks. Very efficiently. So in this case, it's audio generation and then being able to condition on things like text transcripts. The philosophy is like, oh, audio generation as a problem needs to be very efficient, needs to be very real time. And so we need to kind of work on the, the groundwork there to build the sort of the model stack. And then we need to have great training stacks so that we can actually train a model that's high quality that people want to use and, and that, um, actually has a really great experience. So when we, uh, we're putting together the Sonic demo, it was sort of like, uh, we wanted to show that the, Uh, the tech that we were using really can kind of give you something that's really interesting, and text-to-speech is very interesting to me because, uh, you know…
AI assessment note: “how do you kind of build an efficient system that will allow you to do”
Answered raw tape
D 4 · C 4 · P 3 · Cm 3 3.60
Q Like, what's left between here and the ceiling in terms of thinking about the application experience?
A Yeah, I think, like, the way I think about it is, like, would I want to talk to this thing for more than 30 seconds? And if the answer is no, then it's not solved, and if the answer is yes, then it is solved. And I think most text-to-speech systems are not that interesting yet. You don't feel as engaged as you do when you're talking to a human. I, I know there's other, obviously, other reasons you talk to humans, which is, you know, sorry, I don't want to come across as, uh, as crazy here, but yeah, there's a society That we live in, so, um, so we want to talk to people for that reason, obviously, but, but I do think the engagement that you have with these systems is not that high. When you're trying to build these things, you really kind of get so into the weeds on like, oh, I can't say this thing this way, and it's like so boring when it says it that way, and how do I control this part of it to say it like this, you know, the intonation.
AI assessment note: “would I want to talk to this thing for more than 30 seconds?”