Feb 19, 2026 · 1h 22m · mad
Voice AI’s Big Moment: Top Researcher on Why Everything Is Changing (Neil Zeghidour, Gradium AI)
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of The MAD Podcast, host Matt Turck interviews AI researcher Neil Zeghidour about the architectural breakthroughs transforming voice AI, his transition from Google DeepMind to founding Kyutai and Gradium, and the future of real-time full-duplex conversational agents.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 12.1% of the talking time here. How this is scored →
speaking balance: gold is Matt, purple is the guest (3 minute bins)
Neil forcefully dismisses industry standard safety claims, calling audio watermarking an outright scam based on empirical breaks demonstrated in his research.
Hardest push from Matt ▶ 44:40 Challenging voice AI commoditizationMatt directly challenges the guest on whether the subjectivity of voice quality implies that the entire model layer is commoditized commodity tech.
Biggest teaching moment ▶ 57:00 World knowledge density in text versus speechNeil educates the host on why pretraining speech models directly on audio for intelligence is fundamentally flawed compared to text-first architecture.
Matt holds his own ▶ 36:05 Citing Alibaba's Qwen 3 TTS releaseMatt demonstrates deep expertise by citing Alibaba's freshly released Qwen 3 TTS open-source model family to test Gradium's competitive positioning.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Matt as informed peer | Guest teaching | Guest disagreement | Matt pushing back | Why |
|---|---|---|---|---|---|---|
| Episode Highlights: Voice AI's Turning Point | 3 | 2 | 1 | 1 | Matt sets the stage regarding Voice AI's inflection point and asks Neil whether the field is truly at a major moment or still early. Neil collaboratively explains that while latency and accuracy have made huge strides, real-world deployment like noisy environments remains largely unsolved. | |
| Historical Challenges and the Shortage of Voice AI Experts | 3 | 5 | 1 | 1 | Matt asks why voice lagged behind text and vision historically. Neil educates Matt on the historical academic dynamics of speech research, noting that deep learning's first real success was actually speech recognition in 2007, despite speech lacking conference prestige. | |
| Full-Duplex Architecture, Emotional Intelligence, and Future Interfaces | 4 | 5 | 2 | 4 | Matt challenges the premise of ubiquitous voice AI by noting open office privacy constraints where people do not want to talk out loud. Neil reframes the problem, pointing out that voice coding and LLM prompting are shifting paradigms so fast that office layouts themselves may adapt. | |
| Neil Zeghidour's Journey from Applied Math to Audio Language Models | 3 | 6 | 1 | 1 | Neil walks through his background from applied math to Facebook AI, Google Brain, and neural audio codecs. Matt prompts for definitions, and Neil provides a masterclass on how neural codecs compress audio like text, enabling instant zero-shot voice cloning. | |
| Leaving Big Tech and Founding Kyutai AI Lab | 3 | 4 | 1 | 1 | Neil recounts founding Kyutai as an open-science non-profit lab and developing Moshi with a small team in six months. He explains why voice AI does not require 10,000 GPUs, giving small agile teams a distinct advantage over big tech. | |
| Spinning Out Gradium from Open-Source Research Success | 3 | 4 | 2 | 1 | Neil discusses spinning out Gradium for enterprise-grade product needs. He playfully describes his opportunistic research strategy of publishing open papers to frustrate closed-source rivals into joining his team. | |
| Why Focused Startups Can Outmaneuver Big Tech in Voice | 4 | 5 | 2 | 3 | Matt directly asks why big tech giants like OpenAI or Google have not already won voice AI. Neil breaks down parameter efficiency trade-offs in multimodal models and explains why focused developer building blocks beat single generalist assistants. | |
| On-Device Voice AI and the Launch of Pocket TTS | 4 | 5 | 2 | 2 | Matt demonstrates high domain awareness by bringing up Alibaba's newly open-sourced Qwen 3 TTS model family. Neil explains how the industry builds on Moshi's open architecture and details Gradium's Pocket TTS running strictly on CPU. | |
| Rethinking Voice Evaluation Beyond Traditional Benchmarks | 5 | 6 | 3 | 5 | Matt pushes Neil on whether subjective human evaluation implies that voice models are becoming commoditized. Neil strongly rejects this premise, pointing out critical flaws in standard benchmarks like LibriSpeech and highlighting last-mile edge cases. | |
| Cascaded Architecture vs. Direct Speech-to-Speech and Paralinguistics | 3 | 6 | 1 | 1 | Neil explains the technical differences between cascaded architecture (ASR -> LLM -> TTS) and direct speech-to-speech. He details how text bottlenecks discard rich paralinguistic cues like emotion, sarcasm, and tone. | |
| The Problem with Turn-Taking and Rule-Based Voice Activity Detection | 3 | 6 | 4 | 1 | Neil forcefully critiques current turn-taking mechanisms, stating he hates rule-based Voice Activity Detection from the bottom of his heart. He explains how Moshi solved turn-taking using multi-stream token modeling on stereo datasets. | |
| The Hardest Challenge in Voice AI: Noisy and Multi-Speaker Environments | 3 | 6 | 3 | 1 | Neil makes a bold claim, betting that no speech team in the world will solve multi-speaker far-field voice recognition in noisy environments over the next year. He explains the psychoacoustics of binaural sound localization and unconscious human lip reading. | |
| Voice AI Training Data: Quality, Diversity, and Script Generation | 3 | 6 | 1 | 1 | Neil details training data dynamics, explaining why attempting to learn world knowledge from speech audio tokens rather than text is a terrible idea. He shares Gradium's intricate data engineering methods for generating random script prompts without model collapse. | |
| Multilingual Transfer, Low-Resource Languages, and Parallel Datasets | 3 | 5 | 1 | 1 | Neil outlines multilingual transfer learning across language families, noting that 1,000 hours was sufficient for Italian after training on Spanish and Portuguese. He points out that historical datasets like parallel Bible translations remain vital for low-resource languages. | |
| Compute Economics, Hardware Constraints, and On-Demand Compute Scaling | 3 | 5 | 1 | 1 | Neil details compute economics for voice agents, such as gaming NPCs in $70 video games where massive API costs are non-viable. He highlights Gradium's lean 15-person structure working alongside forward deployment partners. | |
| Voice Cloning Capabilities vs. Natural Language Voice Design | 3 | 5 | 1 | 1 | Neil differentiates voice cloning of specific IP/personalities from natural language voice design. He explains why enterprise applications lean heavily toward promptable voice design over exact audio replication. | |
| Discussing Audio Watermarking, Deepfakes, and Voice Cloning Privacy | 4 | 6 | 5 | 2 | Matt raises deepfake and privacy concerns surrounding voice cloning. Neil aggressively dismisses industry watermarking solutions as a scam, citing empirical breaking methods published in the Moshi paper. | |
| Exploring Multimodal AI, Audio-Visual Integration, and Bridger Clone | 4 | 5 | 1 | 1 | Matt prompts a discussion on multimodal integration between voice and video screens. Neil highlights audiovisual speech understanding and mentions consumer experiments like the Bridger Clone app. | |
| Building Gradium in Paris and the French AI Ecosystem | 5 | 5 | 4 | 4 | Matt brings up US tech cynicism toward European AI, citing President Macron's PR misfire regarding AI funding. Neil responds with fierce regional pride, defending Parisian talent density while mocking performative Silicon Valley culture like coding at the gym. |