The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Mati Staniszewski no published score: only 6 usable exchanges on raw tape, and a fair score needs 8+ · coarse estimate ≈4.0/5 from 6 raw tape exchanges record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
6exchanges match
6on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q like, the research and the product effort? Does that make sense? Or like thinking about new markets? And maybe wrapped up in that question too is just like, Well, where are we in quality on, on voice as well? Because if, if I, I would sort of claim like if the models are not good enough for certain use cases at all, like kind of doesn't make sense. Do product?

A And I think that's right. It's, it's almost exactly like when we, when we started originally, what we, what we did was try to actually use existing models that were in the market and kind of optimize them for our first use case was actually starting with combination of, of narration and dubbing, and then on that creative side. And, um, We realized pretty quickly that the models that existed just produced such a robotic and, and not, not good speech that people didn't want to listen to it. And that's where my co-founder's genius came in, where he was able to assemble the team and, and do a lot of the research himself to actually create new version of, of creating that work. But like to your question, I think that the way we are kind of organized internally and how we think about sequencing a lot of that was looking at the first problem. And then creating effectively a lab around that problem, which is like a combination of mighty researchers, engineers, operators to go after that problem. And the first problem was the problem of voice. So how can we recreate the, the voice? And like you say, it needs to have that research expertise to be able to do that well. So we started with effectively a voice lab, which was that mission of, can we narrate the work in, in, in a better way? There was a combination of roughly five people that were That we're doing that work, and then sequence …

AI assessment note: “sequence the research first, and then build a simple layer on top of that”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q imagine for a lot of your customers, it's not like they, like, know how to choose good voice. So how do you, how do you deal with that problem? Like, is it like, hey, I make a clone, and like, that sounds like me, and I believe it, and I'm gonna try all of these different options, or, or, or, You know, actually, are you teaching people to do eval?

A It's a great question, because I think there are, like, two big problems. One is, like, how do you benchmark the general space in audio, where, like you say, it's, like, so dependent on the specific voice, let alone, like, if you are training into interactive, then it's, like, even more tricky. Um, and then the second piece, which is, as you are working on a specific use case, how you select a voice. So I'll take the second one first, which is, uh, we have, like, a voice sommelier, effectively, with us. We work with, with enterprises. We, we, we deploy that person to Work with them and help them navigate. That person is like a voice coach, has an incredible voice themselves, and, uh, and now we have, like, a team under that person that, like, will partner to help you find what's the right branding.

AI assessment note: “we have, like, a voice sommelier, effectively... We deploy that person to Work with them”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q I might get there include working with a Palantir or a large consulting firm, uh, working with 11 or a like platform technology company or, or like an open AI or something, right? Let's talk about that. Uh, or working with a sort of more use case oriented company like Sierra, right? How do you think about how people are making that decision or how they should make that decision?

A The, so, so my past is also in Palantir, so I started exactly kind of from, from that side, and we do blend a lot of the forward deployed engineering inside of the company too. As I think about the kind of our offering and, and the customers making that choice, if you're looking just like, let's say, like one pointed solution, uh, and only that one, then likely we aren't the best choice. If you are looking to deploy that across a plethora of different experiences, so Be it customer support, but then you also want internal training, and you might want to elevate your sales part and actually increase the top line with new experiences of how you engage customers beyond that kind of reactive piece. Um, then it's a great platform to build, and then we effectively, as we engage with customers, combine that platform work with, uh, with our engineering resources to help those companies deploy on that. Or, which we also see increasingly in, um, in Fortune 500, G two, G 2000, where they will want to build Parts of the things themselves, because they already have a lot of the investments in that platform, while then engage us on some of the, the new ones and combine those. And, and, and I think that our model and the way it's different to a lot of the use case specific ones is that our platform is relatively open, where you can use pieces of that platform and not all of them, um, for, for…

AI assessment note: “if you're looking just like, let's say, like one pointed solution... likely we aren't”

Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q and Midjourney and Suno and Hazen in this category, and I think I think there's, like, this overall sense of, like, who really wants to do this? Um, what was your initial read of, like, how many people want to make voices, or what made you believe that was going to be much broader than, you know, Like if I look at dubbing, for example, it's not a huge market.

A I think first piece was, which is, as you mentioned, there is like a very, it's very tricky to do both the product and the research. I'm in a, in a lucky position that I, uh, that my co-founder and I know each other for 15 years. I think he's the smartest person I know and has been able to create a lot of that research work to be able to create that foundation to then elevate that experience. Um, but both of us are from, from Poland originally. And the original belief came from Poland. It's a, it's a very peculiar thing, but if you, if you watch a movie in Polish language, a foreign movie in Polish language, all the voices, whether it's a male voice or a female voice, are narrated with one single character. So you have, like, a flat delivery for everything in a movie.

AI assessment note: “the original belief came from Poland... if you watch a movie in Polish language”

Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q Can you talk about some of the research that you're doing now and then how you think about like the cadence of delivery and what's worth working on?

A We have now a number of, of different initiatives across the audio space and there are, there are kind of two big buckets and, and, and roughly they will relate to that creative and agent side. On the creative side, what this means, uh, we did text to speech models that are controllable. Uh, we then added speech to text model that transcribes in high accurate way, but across a low, uh, resource languages as well. So covering almost a hundred languages. Then created a music model, a fully licensed music model. Um, and as you think about the future, it's how those models will also interact with some of the visual space. So that's, uh, a lot of effort in how you can get the best of audio and then potentially combine that with existing video that you have to, to, to really have the best delivery. And then on the agent side, it's of course how you optimize The real-time speech-to-text, real-time text-to-speech. We just released our speech-to-text model, Scribe v.II, which is under a 150 milliseconds, 93.5% accuracy across the top 30 languages on, on Flourers. And it's only top 30 here because we serve so many others, but most of the people don't. So, uh, so it's, uh, so it's beating, beating the, the, all the models on, on benchmarks. But as you think about the future, it's also the orchestration piece of how you bring speech-to-text LM and text-to-speech. We are releasing, we'll be…

AI assessment note: “We have now a number of, of different initiatives across the audio space”

Answered raw tape D 4 · C 4 · P 3 · Cm 3 3.60

Q And you, you, you can't, you have to very much decide your own narrative at this point in time. I think, correct me if I'm wrong, like in 20, 22 and 23, you probably heard a lot of people say like, Google can do this, and OpenAI can do this, and like, why do you get to persist working on voice anyway as a general capability? What, what's the answer?

A That also adds, adds kind of another element to, to, to that, a couple of the other previous questions where, what is agents work? What is the creative work? Deploy the value in those, in those work, you need a very strong product layer. You need integrations, you need to help people deploy the work, which is the most common piece, but our superpower and our focus for a long time was Building the foundational models to actually make that experience seamless. And as I think about the companies in the market, they will optimize for a lot of other things, and that, that will be, like, the differentiator, um, in our case, where we will make the whole experience, especially with voice, seamless, human, controllable in a, in a much better way.

AI assessment note: “our superpower and our focus for a long time was Building the foundational models”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.