The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Russell D'Sa no published score: only 7 usable exchanges on raw tape, and a fair score needs 8+ · coarse estimate ≈4.5/5 from 7 raw tape exchanges record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
7exchanges match
7on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q information as quickly and condensed as possible. Other people might, might prefer an agent that speaks slower, And really draws things out and gives the full context and then lets them answer and, and kind of ping pong back and forth that way. Uh, is there any movement towards like that level of fine tuning that you think might happen in the future? Is it already happening? I don't know.

A Yeah, it's already happening. So I think that there's kind of two, uh, different flavors of this, um, and you can kind of combine them together, uh, over time. So the first one is that if you've interacted with like the, the real time models, like the ones that are So there's open AI real time API. There's Gemini multimodal live API. Um, there's a few others coming out as well. And, um, for all of these models, uh, that natively understand, uh, audio, um, you can actually tell them to, you know, whisper or slow down or, uh, speed up or act hyper. You can kind of give them an explicit signal of, Um, the style or the way that you want them to communicate with you. So that's already available and possible with these models. Um, then there's another part of it, which I think will come, uh, in the next, you know, year or two, uh, where the model will implicitly be intelligent enough to pick up on what your pacing is or your state of mind is just based on the way that you're talking and expressing yourself. Uh, and also if we weave computer vision into it, It might see you and understand from visuals, okay, this person is stressed, or this person seems like they're in a hurry, or this person is calm and like relaxed, and you can kind of tell that by your body language and by the way you're speaking, and it can automatically adjust you in the same way a human would be able to.

AI assessment note: “Yeah, it's already happening. So I think that there's kind of two”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q if you go back like five years ago, people were really excited about voice, but then you had these sort of like high profile, uh, companies that didn't quite work, but you know, maybe in many ways they were just too early. So What are you seeing on the, on the startup side in terms of new companies being formed specifically to leverage voice and LiveKit and the underlying models?

A Yeah, we're seeing, uh, probably about, um, you know, we're doing like thousands and thousands, many, several thousands of, um, of like signups to the, to the cloud product or commercial product. And, um, most of those, the vast majority are, are startups and growing companies. And out of those, probably around 75% or so, 80% of those signups are, uh, are voice AI companies that are building voice agents. And so the way that I kind of see The market, uh, segmented to a degree is that you have the large AI labs and they have popular, you know, consumer apps. And a lot of them are building kind of open-ended voice agents that you can talk to about kind of anything, right? They're assistants. You can, they do question answering. They do therapy for mental health, all kinds of language learning in the case of speak. Um, and then on the other side, you have these kind of pockets that are what I call kind of voice native systems. And those are really anything you pick up the telephone, uh, To, you know, when you call someone on the other end, call a business, there's someone that answers that line and they're either, you know, doing patient intake at a hospital or they're doing loan qualification or insurance eligibility checking. Um, there's a lot of these kind of business process flows where, um, these are, there are pockets that are like really large in nature. So customer support…

AI assessment note: “probably around 75% or so, 80% of those signups are, uh, are voice AI companies”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q information as quickly and condensed as possible. Other people might, might prefer an agent that speaks slower, And really draws things out and gives the full context and then lets them answer and, and kind of ping pong back and forth that way. Uh, is there any movement towards like that level of fine tuning that you think might happen in the future? Is it already happening? I don't know.

A Yeah, it's already happening. So I think that there's kind of two, uh, different flavors of this, um, and you can kind of combine them together, uh, over time. So the first one is that if you've interacted with like the, the real time models, like the ones that are So there's open AI real time API. There's Gemini multimodal live API. Um, there's a few others coming out as well. And, um, for all of these models, uh, that natively understand, uh, audio, um, you can actually tell them to, you know, whisper or slow down or, uh, speed up or act hyper. You can kind of give them an explicit signal of, Um, the style or the way that you want them to communicate with you. So that's already available and possible with these models. Um, then there's another part of it, which I think will come, uh, in the next, you know, year or two, uh, where the model will implicitly be intelligent enough to pick up on what your pacing is or your state of mind is just based on the way that you're talking and expressing yourself. Uh, and also if we weave computer vision into it, It might see you and understand from visuals, okay, this person is stressed, or this person seems like they're in a hurry, or this person is calm and like relaxed, and you can kind of tell that by your body language, and by the way you're speaking, and it can automatically adjust you in the same way a human would be able to.

AI assessment note: “Yeah, it's already happening. So I think that there's kind of two, uh, different flavors”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Well, anything in voice agents, or just improving the customer experience broadly, like surprise and delight, right?

A Well, you know, it's one of those things that I think is funny, like people talk about hallucinations as this bad thing, but in a lot of ways, for like, kind of this speech-to-speech interaction model, where you're talking to an AI, Hallucination is a feature. So if you remember like your Alexa or Siri, like when you were talking about this, you know, 20, 10, 20, 12, you go and you ask it a question and it's just going through a bunch of like a decision tree. And if it hits, you know, a dead end and it can't answer the question, it just doesn't do anything or it says I can't answer the question. And so the, the second that it, it can't do something you expect it to do all of a sudden you, you know, it's a, it's such a punishing user experience. You just don't want to use the thing anymore and you stop using the thing, but. With kind of these new models that actually, you know, hallucinate answers into existence, even if it doesn't know the answer, it tries, um, you always get a response, uh, that comes back from the model. And for, for like voice interfaces that actually can be more of a feature than a bug, um, it also helps that these models.

AI assessment note: “for like kind of this speech-to-speech interaction model... Hallucination is a feature.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Oh, well, well, anything in voice agents, or just, just, just improving the customer experience broadly, like, surprise and delight, right?

A Well, you know, it's one of those things that I think is funny, like, people talk about hallucinations as this bad thing, um, but in a lot of ways for, like, kind of this speech-to-speech interaction model where you're talking to an AI, Hallucination is a feature. So if you remember like your Alexa or Siri, like when you were talking about this, you know, 2010, 20 12, you go and you ask it a question and it's just going through a bunch of like a decision tree. And if it hits, you know, a dead end and it can't answer the question, it just doesn't do anything or it says I can't answer the question. And so the second that it can't do something you expect it to do, all of a sudden you, you know, it's a, it's such a punishing user experience. You just don't want to use the thing anymore and you stop using the thing. But With kind of these new models that actually, you know, hallucinate answers into existence, even if it doesn't know the answer, it tries. Um, you always get a response, uh, that comes back from the model and for, for like voice interfaces that actually can be more of a feature than a bug. Um, it also helps that these models. Yeah.

AI assessment note: “for like voice interfaces that actually can be more of a feature than a bug”

Answered raw tape D 4 · C 5 · P 4 · Cm 3 4.15

Q Well, anything in voice agents, or just improving the customer experience broadly, like surprise and delight, right?

A Well, you know, it's one of those things that I think is funny, like people talk about hallucinations as this bad thing, but in a lot of ways, for like, kind of this speech-to-speech interaction model, where you're talking to an AI, Hallucination is a feature. So if you remember like your Alexa or Siri, like when you were talking about this, you know, 20, 10, 20, 12, you go and you ask it a question and it's just going through a bunch of like a decision tree. And if it hits, you know, a dead end and it can't answer the question, it just doesn't do anything or it says I can't answer the question. And so the, the second that it, it can't do something you expect it to do all of a sudden you, you know, it's a, it's such a punishing user experience. You just don't want to use the thing anymore and you stop using the thing, but. With kind of these new models that actually, you know, hallucinate answers into existence, even if it doesn't know the answer, it tries, um, you always get a response, uh, that comes back from the model. And for, for like voice interfaces that actually can be more of a feature than a bug, um, it also helps that these models.

AI assessment note: “for like, kind of this speech-to-speech interaction model... Hallucination is a feature”

Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q that once mid-journey gets good enough, uh, or kind of hits some, some peak, that they would just bake it down into silicon, and we would just have image generation models, you know, ASICs, essentially, like what happened with Bitcoin. Uh, is that the future here? Not that you would pivot into hardware, but maybe you would, like, vend your software into a hardware provider at that, at that level?

A Um, well, so we, like, I mean, we partner with, like, folks like Cerebris and Drock. Um, and so we allow you to kind of plug in their models and plug in effectively their hardware accelerated inference. Um, and, and so we're, we're compatible with that world. I think on the inference side, you can definitely, I think it's going to continue to push as low as it can get. Um, there, there's obviously some limits, you know, it's a trade-off between, Uh, kind of capabilities and level of knowledge and how fast you can kind of, kind of, you know, run that, that pass through, uh, the model to get the result out. So there are trade-offs, of course, that follow the laws of physics, but There's also kind of diminishing returns after a while. To give you an example, I once built this, uh, this Cerebris demo. I used like a Lama seven B or eight B, uh, Lama eight B have to remember these numbers, uh, on the primary accounts, but, uh, Lama eight B hooked up to Cerebris. And, uh, I got a bunch of feedback on that voice demo that the model was responding too fast and can you slow it down? And it's kind of going off the rails a little bit.

AI assessment note: “we partner with, like, folks like Cerebris and Drock... we're compatible with that world.”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 500 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.