May 6, 2025 · 20m · latent-space

Voice AI Masterclass — Kwindla Hultman Kramer and swyx

Kwindla Hultman Kramer · 11m spoken Shawn Wang · 6m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this Latent Space episode, Shawn Wang (swyx) and Daily CEO Kwindla Hultman Kramer announce the Voice AI Masterclass, exploring the core architectural pillars, open-source orchestration tools, and hardware applications required to build production-grade conversational AI systems.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 5.8 Guest teaching 4.5 Guest disagreement 1.2 The hosts pushing back 2.0
05100:0010:0020:001:59–4:22 · The hosts as informed peer 5/10 The Three Core Pillars of Voice AI Engineering swyx opens by framing the syllabus and asking for the core curriculum structure, allowing Kwindla to break down the three primary pillars of voice AI engineering. The exchange is deeply collaborative and explanatory as Kwindla details production hurdles, local models, and real-time video.4:22–7:17 · The hosts as informed peer 6/10 Real-Time Interactive Video and Emerging Consumer Interfaces swyx pushes back skeptically on interactive video avatars, claiming they feel robotic and questioning their actual adoption. Kwindla counters directly by noting that Daily powers infrastructure for Tavus and sees massive traction, while swyx demonstrates technical knowledge regarding 3D rigging and video generation constraints.7:17–11:37 · The hosts as informed peer 6/10 Voice Model Architectures and Open Source Ecosystems swyx brings up open source models like Dia and Parakeet and their relation to Soundstorm. Kwindla educates on the differences between experimental research and production-grade transcription models before discussing emerging ecosystem tools like Layercode.11:37–15:30 · The hosts as informed peer 7/10 Solving Core Engineering Challenges: Turn Detection and Parallelism swyx asks precise architectural questions regarding semantic turn detection and whether models can think while simultaneously listening. Kwindla provides technical depth on why state-of-the-art models currently fail at concurrent thinking and how PipeCat orchestrates multi-agent parallel processes to simulate it.15:30–18:37 · The hosts as informed peer 6/10 Voice AI in Production Telephony and Hardware Robotics swyx probes market share in telephony and hardware integration challenges, sharing his own attempts at DIY smart speakers. Kwindla explains the practical dominance of Twilio in production voice AI and showcases physical robotics prototypes.18:37–20:49 · The hosts as informed peer 5/10 Course Community Vision and the Future of Voice Interfaces The conversation concludes on a collaborative and promotional note, highlighting the community vision behind the course and reflecting on the explosive growth of voice engineering tracks.1:59–4:22 · Guest teaching 4/10 The Three Core Pillars of Voice AI Engineering swyx opens by framing the syllabus and asking for the core curriculum structure, allowing Kwindla to break down the three primary pillars of voice AI engineering. The exchange is deeply collaborative and explanatory as Kwindla details production hurdles, local models, and real-time video.4:22–7:17 · Guest teaching 5/10 Real-Time Interactive Video and Emerging Consumer Interfaces swyx pushes back skeptically on interactive video avatars, claiming they feel robotic and questioning their actual adoption. Kwindla counters directly by noting that Daily powers infrastructure for Tavus and sees massive traction, while swyx demonstrates technical knowledge regarding 3D rigging and video generation constraints.7:17–11:37 · Guest teaching 4/10 Voice Model Architectures and Open Source Ecosystems swyx brings up open source models like Dia and Parakeet and their relation to Soundstorm. Kwindla educates on the differences between experimental research and production-grade transcription models before discussing emerging ecosystem tools like Layercode.11:37–15:30 · Guest teaching 6/10 Solving Core Engineering Challenges: Turn Detection and Parallelism swyx asks precise architectural questions regarding semantic turn detection and whether models can think while simultaneously listening. Kwindla provides technical depth on why state-of-the-art models currently fail at concurrent thinking and how PipeCat orchestrates multi-agent parallel processes to simulate it.15:30–18:37 · Guest teaching 5/10 Voice AI in Production Telephony and Hardware Robotics swyx probes market share in telephony and hardware integration challenges, sharing his own attempts at DIY smart speakers. Kwindla explains the practical dominance of Twilio in production voice AI and showcases physical robotics prototypes.18:37–20:49 · Guest teaching 3/10 Course Community Vision and the Future of Voice Interfaces The conversation concludes on a collaborative and promotional note, highlighting the community vision behind the course and reflecting on the explosive growth of voice engineering tracks.1:59–4:22 · Guest disagreement 1/10 The Three Core Pillars of Voice AI Engineering swyx opens by framing the syllabus and asking for the core curriculum structure, allowing Kwindla to break down the three primary pillars of voice AI engineering. The exchange is deeply collaborative and explanatory as Kwindla details production hurdles, local models, and real-time video.4:22–7:17 · Guest disagreement 3/10 Real-Time Interactive Video and Emerging Consumer Interfaces swyx pushes back skeptically on interactive video avatars, claiming they feel robotic and questioning their actual adoption. Kwindla counters directly by noting that Daily powers infrastructure for Tavus and sees massive traction, while swyx demonstrates technical knowledge regarding 3D rigging and video generation constraints.7:17–11:37 · Guest disagreement 1/10 Voice Model Architectures and Open Source Ecosystems swyx brings up open source models like Dia and Parakeet and their relation to Soundstorm. Kwindla educates on the differences between experimental research and production-grade transcription models before discussing emerging ecosystem tools like Layercode.11:37–15:30 · Guest disagreement 1/10 Solving Core Engineering Challenges: Turn Detection and Parallelism swyx asks precise architectural questions regarding semantic turn detection and whether models can think while simultaneously listening. Kwindla provides technical depth on why state-of-the-art models currently fail at concurrent thinking and how PipeCat orchestrates multi-agent parallel processes to simulate it.15:30–18:37 · Guest disagreement 1/10 Voice AI in Production Telephony and Hardware Robotics swyx probes market share in telephony and hardware integration challenges, sharing his own attempts at DIY smart speakers. Kwindla explains the practical dominance of Twilio in production voice AI and showcases physical robotics prototypes.18:37–20:49 · Guest disagreement 0/10 Course Community Vision and the Future of Voice Interfaces The conversation concludes on a collaborative and promotional note, highlighting the community vision behind the course and reflecting on the explosive growth of voice engineering tracks.1:59–4:22 · The hosts pushing back 1/10 The Three Core Pillars of Voice AI Engineering swyx opens by framing the syllabus and asking for the core curriculum structure, allowing Kwindla to break down the three primary pillars of voice AI engineering. The exchange is deeply collaborative and explanatory as Kwindla details production hurdles, local models, and real-time video.4:22–7:17 · The hosts pushing back 4/10 Real-Time Interactive Video and Emerging Consumer Interfaces swyx pushes back skeptically on interactive video avatars, claiming they feel robotic and questioning their actual adoption. Kwindla counters directly by noting that Daily powers infrastructure for Tavus and sees massive traction, while swyx demonstrates technical knowledge regarding 3D rigging and video generation constraints.7:17–11:37 · The hosts pushing back 2/10 Voice Model Architectures and Open Source Ecosystems swyx brings up open source models like Dia and Parakeet and their relation to Soundstorm. Kwindla educates on the differences between experimental research and production-grade transcription models before discussing emerging ecosystem tools like Layercode.11:37–15:30 · The hosts pushing back 2/10 Solving Core Engineering Challenges: Turn Detection and Parallelism swyx asks precise architectural questions regarding semantic turn detection and whether models can think while simultaneously listening. Kwindla provides technical depth on why state-of-the-art models currently fail at concurrent thinking and how PipeCat orchestrates multi-agent parallel processes to simulate it.15:30–18:37 · The hosts pushing back 2/10 Voice AI in Production Telephony and Hardware Robotics swyx probes market share in telephony and hardware integration challenges, sharing his own attempts at DIY smart speakers. Kwindla explains the practical dominance of Twilio in production voice AI and showcases physical robotics prototypes.18:37–20:49 · The hosts pushing back 1/10 Course Community Vision and the Future of Voice Interfaces The conversation concludes on a collaborative and promotional note, highlighting the community vision behind the course and reflecting on the explosive growth of voice engineering tracks.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 5:05 Kwindla refuting skepticism on avatar traction

Kwindla directly disputes swyx's claim that video avatars lack real demand by citing Daily's internal network metrics powering platforms like Tavus.

Hardest push from the hosts ▶ 4:55 swyx challenging interactive video utility

swyx explicitly challenges the premise that interactive real-time video avatars have traction, calling the experience unnatural and robotic.

Biggest teaching moment ▶ 7:49 Kwindla contrasting Dia and Parakeet models

Kwindla clarifies the open voice landscape by categorizing Dia as an unhinged dynamic student experiment and Nvidia's Parakeet as an enterprise-grade transcription model.

The host holds their own ▶ 6:39 swyx breaking down avatar animation technical limits

swyx showcases deep domain knowledge by detailing how real-time avatar animation relies on constrained bone-structure rigging rather than general video generation.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
The Three Core Pillars of Voice AI Engineering 5411 swyx opens by framing the syllabus and asking for the core curriculum structure, allowing Kwindla to break down the three primary pillars of voice AI engineering. The exchange is deeply collaborative and explanatory as Kwindla details production hurdles, local models, and real-time video.
Real-Time Interactive Video and Emerging Consumer Interfaces 6534 swyx pushes back skeptically on interactive video avatars, claiming they feel robotic and questioning their actual adoption. Kwindla counters directly by noting that Daily powers infrastructure for Tavus and sees massive traction, while swyx demonstrates technical knowledge regarding 3D rigging and video generation constraints.
Voice Model Architectures and Open Source Ecosystems 6412 swyx brings up open source models like Dia and Parakeet and their relation to Soundstorm. Kwindla educates on the differences between experimental research and production-grade transcription models before discussing emerging ecosystem tools like Layercode.
Solving Core Engineering Challenges: Turn Detection and Parallelism 7612 swyx asks precise architectural questions regarding semantic turn detection and whether models can think while simultaneously listening. Kwindla provides technical depth on why state-of-the-art models currently fail at concurrent thinking and how PipeCat orchestrates multi-agent parallel processes to simulate it.
Voice AI in Production Telephony and Hardware Robotics 6512 swyx probes market share in telephony and hardware integration challenges, sharing his own attempts at DIY smart speakers. Kwindla explains the practical dominance of Twilio in production voice AI and showcases physical robotics prototypes.
Course Community Vision and the Future of Voice Interfaces 5301 The conversation concludes on a collaborative and promotional note, highlighting the community vision behind the course and reflecting on the explosive growth of voice engineering tracks.

Statements from this episode (11)

Insight
Multimodal Real-Time Agents Require an Entirely New Programming Paradigm
“And it turns out that if you're building like agents that are multi-modal, multi-turn, real-time, it's just a totally different shape of programming problems and best practices than most other kinds of AI development even.”
Kwindla Hultman Kramer May 6, 2025 ▶ 2:12
Assertion Not checkable as stated
Speech-to-Speech Models Are Not Yet Widely Used in Production
“What's happening with speech to speech models and APIs, super hot topic, not widely used yet in production.”
Kwindla Hultman Kramer May 6, 2025 ▶ 3:38
Prediction Not checkable as stated
Real-Time Video AI Will Inflect by Early 2026
“Like I really think real-time video is going to hit the same inflection point that voice did by the end of the year or early next year.”
Kwindla Hultman Kramer May 6, 2025 ▶ 4:03
Assertion Not checkable as stated
Enterprise B2B Unexpectedly Drove the Voice AI Monetization Inflection Point
“A little bit to everybody's surprise, the voice AI monetization pull actually came from enterprise and B to B use cases. So a little bit, everybody surprised the inflection point with voice AI in terms of monetizable use cases came on the business side. It's l…”
Kwindla Hultman Kramer May 6, 2025 ▶ 5:35
Prediction Not checkable as stated
Real-Time Interactive AI Video Will Hit Consumers Before Enterprise
“I think real-time video may actually hit on the consumer side first. When it, when it's right, when it's on the right side of the uncanny valley, it's really, really compelling. And I think we're just starting to see some of that.”
Kwindla Hultman Kramer May 6, 2025 ▶ 6:01
Prediction Not checkable as stated
The Next TikTok Will Focus on Hyper-Personalized Interactive AI Video
“Well, I think we're going to have friends that are video in all our group chats, and the next TikTok is going to be not just hyper-personalized recorded content, but hyper-personalized interactive content.”
Kwindla Hultman Kramer May 6, 2025 ▶ 6:27
Assertion Supported
NVIDIA's 600M Parameter Parakeet Model Tops Speech Transcription Leaderboards
“And then Parakeet is Nvidia's new speech model, speech transcription model. That's number one on the leaderboards. And it's just like very enterprise tuned, like really, really rock solid, reliable at fairly small number of weights, like six hundred million pa…”
Kwindla Hultman Kramer May 6, 2025 ▶ 8:07
Assertion Supported
No State-of-the-Art Model Can Reason While Processing Continuous Input
“None of the SOTA models can kind of think while also taking input.”
Kwindla Hultman Kramer May 6, 2025 ▶ 14:18
Assertion Supported
Kyutai's Bidirectional Moshi Architecture Has Not Scaled to Large LLMs
“The, that Kyutai Moshi model you mentioned, which was my favorite academic paper last year, is a step towards a truly bi-directional streaming in both directions, thinking all the time, LLM. That work has not been kind of scaled up to, you know, large LLM size…”
Kwindla Hultman Kramer May 6, 2025 ▶ 14:25
Assertion Not checkable as stated
99% of Current Monetizable Voice AI Use Cases Are Telephony
“99% of the monetizable voice AI use cases today are telephony.”
Kwindla Hultman Kramer May 6, 2025 ▶ 15:48
Prediction Not checkable as stated
Up to 75% of Future UX Interfaces Will Be Voice-Driven
“UX is going to be, you know, 50%, 60%, 75% voice in the future. I a hundred percent believe that, and I did not believe that, you know, two years ago, but the trend line is just really, I think, clear.”
Kwindla Hultman Kramer May 6, 2025 ▶ 20:24
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.