Apr 26, 2025 · 17m · tbpn

Why Hallucination Is Good For AI Voice Agents | Russell D'Sa on TBPN

Russell D'Sa · 11m spoken John Coogan · 2m spoken Jordi Hays · 2m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this interview, LiveKit co-founder and CEO Russell D'Sa explores the technical breakthroughs in latency and large language models enabling real-time voice AI, while discussing developer adoption trends and LiveKit's $45 million Series B funding milestone.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 31.8% of the talking time here. How this is scored →

The hosts as informed peer 5.3 Guest teaching 4.5 Guest disagreement 1.5 The hosts pushing back 1.8
05100:0010:002:13–6:00 · The hosts as informed peer 6/10 Why AI Hallucinations Benefit Voice Interface User Experience Coogan leverages his past experience working in voice tech at Vlingo/Dragon to ask why speech recognition is still flawed. D'Sa delivers a contrarian reframe, explaining why hallucination serves as a critical UX feature rather than a bug in conversational voice interfaces.6:01–9:49 · The hosts as informed peer 7/10 Hardware Acceleration, Diminishing Returns, and Adaptive Agents Coogan demonstrates substantial technical knowledge by drawing comparisons to Etched and Bitcoin ASICs, inquiring about silicon-level inference and runtime adaptive pacing. D'Sa elaborates on diminishing returns in latency and the emergence of multimodal pacing detection.9:50–13:39 · The hosts as informed peer 4/10 The Startup Landscape and Enterprise Customer Service Automation Hays probes the startup ecosystem and adoption rates across enterprise verticals. D'Sa breaks down the market distinction between broad consumer AI labs and narrow voice-native business process automation.13:40–16:10 · The hosts as informed peer 4/10 Real-Time Synchronization: The Playback Live Sports Case Study Coogan brings up a previous podcast guest's company, Playback, prompting D'Sa to explain the technical mechanics of broadcast desynchronization versus LiveKit's synchronized live video distribution.2:13–6:00 · Guest teaching 5/10 Why AI Hallucinations Benefit Voice Interface User Experience Coogan leverages his past experience working in voice tech at Vlingo/Dragon to ask why speech recognition is still flawed. D'Sa delivers a contrarian reframe, explaining why hallucination serves as a critical UX feature rather than a bug in conversational voice interfaces.6:01–9:49 · Guest teaching 4/10 Hardware Acceleration, Diminishing Returns, and Adaptive Agents Coogan demonstrates substantial technical knowledge by drawing comparisons to Etched and Bitcoin ASICs, inquiring about silicon-level inference and runtime adaptive pacing. D'Sa elaborates on diminishing returns in latency and the emergence of multimodal pacing detection.9:50–13:39 · Guest teaching 4/10 The Startup Landscape and Enterprise Customer Service Automation Hays probes the startup ecosystem and adoption rates across enterprise verticals. D'Sa breaks down the market distinction between broad consumer AI labs and narrow voice-native business process automation.13:40–16:10 · Guest teaching 5/10 Real-Time Synchronization: The Playback Live Sports Case Study Coogan brings up a previous podcast guest's company, Playback, prompting D'Sa to explain the technical mechanics of broadcast desynchronization versus LiveKit's synchronized live video distribution.2:13–6:00 · Guest disagreement 3/10 Why AI Hallucinations Benefit Voice Interface User Experience Coogan leverages his past experience working in voice tech at Vlingo/Dragon to ask why speech recognition is still flawed. D'Sa delivers a contrarian reframe, explaining why hallucination serves as a critical UX feature rather than a bug in conversational voice interfaces.6:01–9:49 · Guest disagreement 1/10 Hardware Acceleration, Diminishing Returns, and Adaptive Agents Coogan demonstrates substantial technical knowledge by drawing comparisons to Etched and Bitcoin ASICs, inquiring about silicon-level inference and runtime adaptive pacing. D'Sa elaborates on diminishing returns in latency and the emergence of multimodal pacing detection.9:50–13:39 · Guest disagreement 1/10 The Startup Landscape and Enterprise Customer Service Automation Hays probes the startup ecosystem and adoption rates across enterprise verticals. D'Sa breaks down the market distinction between broad consumer AI labs and narrow voice-native business process automation.13:40–16:10 · Guest disagreement 1/10 Real-Time Synchronization: The Playback Live Sports Case Study Coogan brings up a previous podcast guest's company, Playback, prompting D'Sa to explain the technical mechanics of broadcast desynchronization versus LiveKit's synchronized live video distribution.2:13–6:00 · The hosts pushing back 3/10 Why AI Hallucinations Benefit Voice Interface User Experience Coogan leverages his past experience working in voice tech at Vlingo/Dragon to ask why speech recognition is still flawed. D'Sa delivers a contrarian reframe, explaining why hallucination serves as a critical UX feature rather than a bug in conversational voice interfaces.6:01–9:49 · The hosts pushing back 2/10 Hardware Acceleration, Diminishing Returns, and Adaptive Agents Coogan demonstrates substantial technical knowledge by drawing comparisons to Etched and Bitcoin ASICs, inquiring about silicon-level inference and runtime adaptive pacing. D'Sa elaborates on diminishing returns in latency and the emergence of multimodal pacing detection.9:50–13:39 · The hosts pushing back 1/10 The Startup Landscape and Enterprise Customer Service Automation Hays probes the startup ecosystem and adoption rates across enterprise verticals. D'Sa breaks down the market distinction between broad consumer AI labs and narrow voice-native business process automation.13:40–16:10 · The hosts pushing back 1/10 Real-Time Synchronization: The Playback Live Sports Case Study Coogan brings up a previous podcast guest's company, Playback, prompting D'Sa to explain the technical mechanics of broadcast desynchronization versus LiveKit's synchronized live video distribution.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 36.5% · guest 63.5%0:00 · the hosts 36.5% · guest 63.5%3:00 · the hosts 23.2% · guest 76.8%3:00 · the hosts 23.2% · guest 76.8%6:00 · the hosts 43.6% · guest 56.4%6:00 · the hosts 43.6% · guest 56.4%9:00 · the hosts 23.7% · guest 76.3%9:00 · the hosts 23.7% · guest 76.3%12:00 · the hosts 32.3% · guest 67.7%12:00 · the hosts 32.3% · guest 67.7%15:00 · the hosts 31.8% · guest 68.2%15:00 · the hosts 31.8% · guest 68.2%
Sharpest disagreement ▶ 3:31 Hallucination Reframed as a Core Feature

D'Sa rejects the conventional industry framing that model hallucinations are purely negative, arguing they are necessary for conversational voice UX.

Hardest push from the hosts ▶ 2:13 Challenging Speech-to-Text Progress

Coogan pushes back on modern voice progress by demanding D'Sa steelman why speech-to-text dictation on smartphones remains frustratingly imperfect.

Biggest teaching moment ▶ 14:30 Explaining Broadcast Desync vs True Real-Time

D'Sa educates the hosts on broadcast architecture, detailing how modern streaming lags 30-60 seconds compared to Apollo-era synchronization.

The host holds their own ▶ 6:01 Coogan Cites Custom Silicon and Hardware ASICs

Coogan displays deep industry fluency by referencing transformer ASICs from Etched and comparing future model baking to Bitcoin mining hardware.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Why AI Hallucinations Benefit Voice Interface User Experience 6533 Coogan leverages his past experience working in voice tech at Vlingo/Dragon to ask why speech recognition is still flawed. D'Sa delivers a contrarian reframe, explaining why hallucination serves as a critical UX feature rather than a bug in conversational voice interfaces.
Hardware Acceleration, Diminishing Returns, and Adaptive Agents 7412 Coogan demonstrates substantial technical knowledge by drawing comparisons to Etched and Bitcoin ASICs, inquiring about silicon-level inference and runtime adaptive pacing. D'Sa elaborates on diminishing returns in latency and the emergence of multimodal pacing detection.
The Startup Landscape and Enterprise Customer Service Automation 4411 Hays probes the startup ecosystem and adoption rates across enterprise verticals. D'Sa breaks down the market distinction between broad consumer AI labs and narrow voice-native business process automation.
Real-Time Synchronization: The Playback Live Sports Case Study 4511 Coogan brings up a previous podcast guest's company, Playback, prompting D'Sa to explain the technical mechanics of broadcast desynchronization versus LiveKit's synchronized live video distribution.

Statements from this episode (10)

Assertion Supported
D'Sa: LiveKit partnered with OpenAI to build ChatGPT voice mode
“We built a demo when ChatGPT first came out, ah, where you could talk to it instead of text with it using LiveKit's infrastructure, and a few months later, OpenAI ended up finding it, ah, wanted to build a voice interface to ChatGPT, And we started to work pre…”
Russell D'Sa Apr 26, 2025 ▶ 1:14
Insight
LiveKit CEO: Hallucination is a feature in voice AI interfaces
“Well, you know, it's one of those things that I think is funny, like, people talk about hallucinations as this bad thing but in a lot of ways for, like, kind of this speech-to-speech interaction model where you're talking to an AI, Hallucination is a feature.”
Russell D'Sa Apr 26, 2025 ▶ 3:31
Assertion Supported
D'Sa: LLM inference is now faster than TTS speech generation
“And now like LLM inference can actually be done in less time than generating speech with TTS.”
Russell D'Sa Apr 26, 2025 ▶ 5:38
Insight
D'Sa: 300-500ms latency crosses the voice AI uncanny valley
“I think, like, getting that latency end to end down, you know, to 300 milliseconds or 500 milliseconds on average for, like, turn latency that, that, that kind of helps you cross this uncanny valley for voice AI.”
Russell D'Sa Apr 26, 2025 ▶ 5:47
Disclosure
D'Sa: LiveKit partners with Cerebras and Groq for hardware inference
“Well, so we, like, I mean, we partner with, like, folks like Cerebris and Drock. And so we allow you to kind of plug in their models and plug in effectively their hardware accelerated inference. And so we're compatible with that world.”
Russell D'Sa Apr 26, 2025 ▶ 6:36
Insight
D'Sa: Ultra-fast voice AI latency yields diminishing returns
“There's also kind of diminishing returns after a while. To give you an example, I once built this Cerebris demo. I used like a Lama seven B or eight B Lama eight B have to remember these numbers on the primary accounts, but Lama eight B hooked up to Cerebris. …”
Russell D'Sa Apr 26, 2025 ▶ 7:15
Prediction Held up
D'Sa: Voice models will adapt to user pacing within two years
“Then there's another part of it, which I think will come in the next, you know, year or two where the model will implicitly be intelligent enough to pick up on what your pacing is or your state of mind is just based on the way that you're talking and expressin…”
Russell D'Sa Apr 26, 2025 ▶ 9:12
Disclosure
D'Sa: 75-80% of LiveKit cloud signups are building voice AI agents
“We're doing like thousands and thousands, many, several thousands of like signups to the cloud product or commercial product. And most of those, the vast majority are startups and growing companies. And out of those, probably around 75% or so, 80% of those sig…”
Russell D'Sa Apr 26, 2025 ▶ 10:33
Prediction Didn’t hold up
Playback will integrate AI voice commentary overlays for fans
“Playback they're also going to integrate you know, AI into that flow as well, like a voice based commentator or whatever that can watch the game and, Provide an overlay for fans”
Russell D'Sa Apr 26, 2025 ▶ 14:01
Assertion Supported
LiveKit closes $45M Series B led by Altimeter Capital
“And so we closed a series B today. That's led by Altimeter Capital.”
Russell D'Sa Apr 26, 2025 ▶ 16:36
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 500 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.