Russell D'Sa, CEO of LiveKit, predicts voice AI will automatically adjust pacing to match user stress levels and conversational speed.
Insight
LiveKit CEO: Hallucination is a feature in voice AI interfaces
“Well, you know, it's one of those things that I think is funny, like, people talk about hallucinations as this bad thing but in a lot of ways for, like, kind of this speech-to-speech interaction model where you're talking to an AI, Hallucination is a feature.”
Insight
Russ d'Sa: AI hallucination is a feature in voice interfaces
“People talk about hallucinations as this bad thing, but in a lot of ways, for like, kind of this speech-to-speech interaction model, where you're talking to an AI, Hallucination is a feature.”
Assertion Supported
Russ d'Sa: LLM inference is now faster than text-to-speech generation
“Now like LLM inference can actually be done in less time than generating speech with TTS.”
Insight
Russ d'Sa: 300 to 500 ms turn latency crosses voice AI uncanny valley
“Getting that latency end to end down, you know, to 300 milliseconds or 500 milliseconds on average for, like, turn latency that, that, that kind of helps you cross this uncanny valley for voice AI.”
Assertion Supported
LiveKit's d'Sa: LiveKit built the infrastructure for OpenAI's ChatGPT voice modes
“Fast forward, like, about a year and a half, two years, and, ah, we built a demo when ChatGPT first came out, ah, where you could talk to it instead of text with it using LiveKit's infrastructure, and a few months later, OpenAI ended up finding it, ah, wanted …”
Prediction Held up
Russ d'Sa: AI models will implicitly match user pacing within two years
“Then there's another part of it, which I think will come in the next, you know, year or two where the model will implicitly be intelligent enough to pick up on what your pacing is or your state of mind is just based on the way that you're talking and expressin…”