Feb 19, 2026 · 1h 22m · mad

Voice AI’s Big Moment: Top Researcher on Why Everything Is Changing (Neil Zeghidour, Gradium AI)

Neil Zeghidour · 1h 7m spoken Matt Turck · 9m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of The MAD Podcast, host Matt Turck interviews AI researcher Neil Zeghidour about the architectural breakthroughs transforming voice AI, his transition from Google DeepMind to founding Kyutai and Gradium, and the future of real-time full-duplex conversational agents.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 12.1% of the talking time here. How this is scored →

Matt as informed peer 3.5 Guest teaching 5.1 Guest disagreement 1.9 Matt pushing back 1.7
05100:0020:0040:001:00:001:20:000:00–3:27 · Matt as informed peer 3/10 Episode Highlights: Voice AI's Turning Point Matt sets the stage regarding Voice AI's inflection point and asks Neil whether the field is truly at a major moment or still early. Neil collaboratively explains that while latency and accuracy have made huge strides, real-world deployment like noisy environments remains largely unsolved.3:27–7:39 · Matt as informed peer 3/10 Historical Challenges and the Shortage of Voice AI Experts Matt asks why voice lagged behind text and vision historically. Neil educates Matt on the historical academic dynamics of speech research, noting that deep learning's first real success was actually speech recognition in 2007, despite speech lacking conference prestige.7:39–12:55 · Matt as informed peer 4/10 Full-Duplex Architecture, Emotional Intelligence, and Future Interfaces Matt challenges the premise of ubiquitous voice AI by noting open office privacy constraints where people do not want to talk out loud. Neil reframes the problem, pointing out that voice coding and LLM prompting are shifting paradigms so fast that office layouts themselves may adapt.12:55–22:20 · Matt as informed peer 3/10 Neil Zeghidour's Journey from Applied Math to Audio Language Models Neil walks through his background from applied math to Facebook AI, Google Brain, and neural audio codecs. Matt prompts for definitions, and Neil provides a masterclass on how neural codecs compress audio like text, enabling instant zero-shot voice cloning.22:20–27:09 · Matt as informed peer 3/10 Leaving Big Tech and Founding Kyutai AI Lab Neil recounts founding Kyutai as an open-science non-profit lab and developing Moshi with a small team in six months. He explains why voice AI does not require 10,000 GPUs, giving small agile teams a distinct advantage over big tech.27:09–31:26 · Matt as informed peer 3/10 Spinning Out Gradium from Open-Source Research Success Neil discusses spinning out Gradium for enterprise-grade product needs. He playfully describes his opportunistic research strategy of publishing open papers to frustrate closed-source rivals into joining his team.31:26–34:01 · Matt as informed peer 4/10 Why Focused Startups Can Outmaneuver Big Tech in Voice Matt directly asks why big tech giants like OpenAI or Google have not already won voice AI. Neil breaks down parameter efficiency trade-offs in multimodal models and explains why focused developer building blocks beat single generalist assistants.34:01–39:09 · Matt as informed peer 4/10 On-Device Voice AI and the Launch of Pocket TTS Matt demonstrates high domain awareness by bringing up Alibaba's newly open-sourced Qwen 3 TTS model family. Neil explains how the industry builds on Moshi's open architecture and details Gradium's Pocket TTS running strictly on CPU.39:09–46:50 · Matt as informed peer 5/10 Rethinking Voice Evaluation Beyond Traditional Benchmarks Matt pushes Neil on whether subjective human evaluation implies that voice models are becoming commoditized. Neil strongly rejects this premise, pointing out critical flaws in standard benchmarks like LibriSpeech and highlighting last-mile edge cases.46:50–49:39 · Matt as informed peer 3/10 Cascaded Architecture vs. Direct Speech-to-Speech and Paralinguistics Neil explains the technical differences between cascaded architecture (ASR -> LLM -> TTS) and direct speech-to-speech. He details how text bottlenecks discard rich paralinguistic cues like emotion, sarcasm, and tone.49:39–53:29 · Matt as informed peer 3/10 The Problem with Turn-Taking and Rule-Based Voice Activity Detection Neil forcefully critiques current turn-taking mechanisms, stating he hates rule-based Voice Activity Detection from the bottom of his heart. He explains how Moshi solved turn-taking using multi-stream token modeling on stereo datasets.53:29–56:14 · Matt as informed peer 3/10 The Hardest Challenge in Voice AI: Noisy and Multi-Speaker Environments Neil makes a bold claim, betting that no speech team in the world will solve multi-speaker far-field voice recognition in noisy environments over the next year. He explains the psychoacoustics of binaural sound localization and unconscious human lip reading.56:14–1:00:50 · Matt as informed peer 3/10 Voice AI Training Data: Quality, Diversity, and Script Generation Neil details training data dynamics, explaining why attempting to learn world knowledge from speech audio tokens rather than text is a terrible idea. He shares Gradium's intricate data engineering methods for generating random script prompts without model collapse.1:00:50–1:02:55 · Matt as informed peer 3/10 Multilingual Transfer, Low-Resource Languages, and Parallel Datasets Neil outlines multilingual transfer learning across language families, noting that 1,000 hours was sufficient for Italian after training on Spanish and Portuguese. He points out that historical datasets like parallel Bible translations remain vital for low-resource languages.1:02:55–1:06:30 · Matt as informed peer 3/10 Compute Economics, Hardware Constraints, and On-Demand Compute Scaling Neil details compute economics for voice agents, such as gaming NPCs in $70 video games where massive API costs are non-viable. He highlights Gradium's lean 15-person structure working alongside forward deployment partners.1:06:30–1:09:02 · Matt as informed peer 3/10 Voice Cloning Capabilities vs. Natural Language Voice Design Neil differentiates voice cloning of specific IP/personalities from natural language voice design. He explains why enterprise applications lean heavily toward promptable voice design over exact audio replication.1:09:02–1:12:30 · Matt as informed peer 4/10 Discussing Audio Watermarking, Deepfakes, and Voice Cloning Privacy Matt raises deepfake and privacy concerns surrounding voice cloning. Neil aggressively dismisses industry watermarking solutions as a scam, citing empirical breaking methods published in the Moshi paper.1:12:30–1:16:29 · Matt as informed peer 4/10 Exploring Multimodal AI, Audio-Visual Integration, and Bridger Clone Matt prompts a discussion on multimodal integration between voice and video screens. Neil highlights audiovisual speech understanding and mentions consumer experiments like the Bridger Clone app.1:16:29–1:22:28 · Matt as informed peer 5/10 Building Gradium in Paris and the French AI Ecosystem Matt brings up US tech cynicism toward European AI, citing President Macron's PR misfire regarding AI funding. Neil responds with fierce regional pride, defending Parisian talent density while mocking performative Silicon Valley culture like coding at the gym.0:00–3:27 · Guest teaching 2/10 Episode Highlights: Voice AI's Turning Point Matt sets the stage regarding Voice AI's inflection point and asks Neil whether the field is truly at a major moment or still early. Neil collaboratively explains that while latency and accuracy have made huge strides, real-world deployment like noisy environments remains largely unsolved.3:27–7:39 · Guest teaching 5/10 Historical Challenges and the Shortage of Voice AI Experts Matt asks why voice lagged behind text and vision historically. Neil educates Matt on the historical academic dynamics of speech research, noting that deep learning's first real success was actually speech recognition in 2007, despite speech lacking conference prestige.7:39–12:55 · Guest teaching 5/10 Full-Duplex Architecture, Emotional Intelligence, and Future Interfaces Matt challenges the premise of ubiquitous voice AI by noting open office privacy constraints where people do not want to talk out loud. Neil reframes the problem, pointing out that voice coding and LLM prompting are shifting paradigms so fast that office layouts themselves may adapt.12:55–22:20 · Guest teaching 6/10 Neil Zeghidour's Journey from Applied Math to Audio Language Models Neil walks through his background from applied math to Facebook AI, Google Brain, and neural audio codecs. Matt prompts for definitions, and Neil provides a masterclass on how neural codecs compress audio like text, enabling instant zero-shot voice cloning.22:20–27:09 · Guest teaching 4/10 Leaving Big Tech and Founding Kyutai AI Lab Neil recounts founding Kyutai as an open-science non-profit lab and developing Moshi with a small team in six months. He explains why voice AI does not require 10,000 GPUs, giving small agile teams a distinct advantage over big tech.27:09–31:26 · Guest teaching 4/10 Spinning Out Gradium from Open-Source Research Success Neil discusses spinning out Gradium for enterprise-grade product needs. He playfully describes his opportunistic research strategy of publishing open papers to frustrate closed-source rivals into joining his team.31:26–34:01 · Guest teaching 5/10 Why Focused Startups Can Outmaneuver Big Tech in Voice Matt directly asks why big tech giants like OpenAI or Google have not already won voice AI. Neil breaks down parameter efficiency trade-offs in multimodal models and explains why focused developer building blocks beat single generalist assistants.34:01–39:09 · Guest teaching 5/10 On-Device Voice AI and the Launch of Pocket TTS Matt demonstrates high domain awareness by bringing up Alibaba's newly open-sourced Qwen 3 TTS model family. Neil explains how the industry builds on Moshi's open architecture and details Gradium's Pocket TTS running strictly on CPU.39:09–46:50 · Guest teaching 6/10 Rethinking Voice Evaluation Beyond Traditional Benchmarks Matt pushes Neil on whether subjective human evaluation implies that voice models are becoming commoditized. Neil strongly rejects this premise, pointing out critical flaws in standard benchmarks like LibriSpeech and highlighting last-mile edge cases.46:50–49:39 · Guest teaching 6/10 Cascaded Architecture vs. Direct Speech-to-Speech and Paralinguistics Neil explains the technical differences between cascaded architecture (ASR -> LLM -> TTS) and direct speech-to-speech. He details how text bottlenecks discard rich paralinguistic cues like emotion, sarcasm, and tone.49:39–53:29 · Guest teaching 6/10 The Problem with Turn-Taking and Rule-Based Voice Activity Detection Neil forcefully critiques current turn-taking mechanisms, stating he hates rule-based Voice Activity Detection from the bottom of his heart. He explains how Moshi solved turn-taking using multi-stream token modeling on stereo datasets.53:29–56:14 · Guest teaching 6/10 The Hardest Challenge in Voice AI: Noisy and Multi-Speaker Environments Neil makes a bold claim, betting that no speech team in the world will solve multi-speaker far-field voice recognition in noisy environments over the next year. He explains the psychoacoustics of binaural sound localization and unconscious human lip reading.56:14–1:00:50 · Guest teaching 6/10 Voice AI Training Data: Quality, Diversity, and Script Generation Neil details training data dynamics, explaining why attempting to learn world knowledge from speech audio tokens rather than text is a terrible idea. He shares Gradium's intricate data engineering methods for generating random script prompts without model collapse.1:00:50–1:02:55 · Guest teaching 5/10 Multilingual Transfer, Low-Resource Languages, and Parallel Datasets Neil outlines multilingual transfer learning across language families, noting that 1,000 hours was sufficient for Italian after training on Spanish and Portuguese. He points out that historical datasets like parallel Bible translations remain vital for low-resource languages.1:02:55–1:06:30 · Guest teaching 5/10 Compute Economics, Hardware Constraints, and On-Demand Compute Scaling Neil details compute economics for voice agents, such as gaming NPCs in $70 video games where massive API costs are non-viable. He highlights Gradium's lean 15-person structure working alongside forward deployment partners.1:06:30–1:09:02 · Guest teaching 5/10 Voice Cloning Capabilities vs. Natural Language Voice Design Neil differentiates voice cloning of specific IP/personalities from natural language voice design. He explains why enterprise applications lean heavily toward promptable voice design over exact audio replication.1:09:02–1:12:30 · Guest teaching 6/10 Discussing Audio Watermarking, Deepfakes, and Voice Cloning Privacy Matt raises deepfake and privacy concerns surrounding voice cloning. Neil aggressively dismisses industry watermarking solutions as a scam, citing empirical breaking methods published in the Moshi paper.1:12:30–1:16:29 · Guest teaching 5/10 Exploring Multimodal AI, Audio-Visual Integration, and Bridger Clone Matt prompts a discussion on multimodal integration between voice and video screens. Neil highlights audiovisual speech understanding and mentions consumer experiments like the Bridger Clone app.1:16:29–1:22:28 · Guest teaching 5/10 Building Gradium in Paris and the French AI Ecosystem Matt brings up US tech cynicism toward European AI, citing President Macron's PR misfire regarding AI funding. Neil responds with fierce regional pride, defending Parisian talent density while mocking performative Silicon Valley culture like coding at the gym.0:00–3:27 · Guest disagreement 1/10 Episode Highlights: Voice AI's Turning Point Matt sets the stage regarding Voice AI's inflection point and asks Neil whether the field is truly at a major moment or still early. Neil collaboratively explains that while latency and accuracy have made huge strides, real-world deployment like noisy environments remains largely unsolved.3:27–7:39 · Guest disagreement 1/10 Historical Challenges and the Shortage of Voice AI Experts Matt asks why voice lagged behind text and vision historically. Neil educates Matt on the historical academic dynamics of speech research, noting that deep learning's first real success was actually speech recognition in 2007, despite speech lacking conference prestige.7:39–12:55 · Guest disagreement 2/10 Full-Duplex Architecture, Emotional Intelligence, and Future Interfaces Matt challenges the premise of ubiquitous voice AI by noting open office privacy constraints where people do not want to talk out loud. Neil reframes the problem, pointing out that voice coding and LLM prompting are shifting paradigms so fast that office layouts themselves may adapt.12:55–22:20 · Guest disagreement 1/10 Neil Zeghidour's Journey from Applied Math to Audio Language Models Neil walks through his background from applied math to Facebook AI, Google Brain, and neural audio codecs. Matt prompts for definitions, and Neil provides a masterclass on how neural codecs compress audio like text, enabling instant zero-shot voice cloning.22:20–27:09 · Guest disagreement 1/10 Leaving Big Tech and Founding Kyutai AI Lab Neil recounts founding Kyutai as an open-science non-profit lab and developing Moshi with a small team in six months. He explains why voice AI does not require 10,000 GPUs, giving small agile teams a distinct advantage over big tech.27:09–31:26 · Guest disagreement 2/10 Spinning Out Gradium from Open-Source Research Success Neil discusses spinning out Gradium for enterprise-grade product needs. He playfully describes his opportunistic research strategy of publishing open papers to frustrate closed-source rivals into joining his team.31:26–34:01 · Guest disagreement 2/10 Why Focused Startups Can Outmaneuver Big Tech in Voice Matt directly asks why big tech giants like OpenAI or Google have not already won voice AI. Neil breaks down parameter efficiency trade-offs in multimodal models and explains why focused developer building blocks beat single generalist assistants.34:01–39:09 · Guest disagreement 2/10 On-Device Voice AI and the Launch of Pocket TTS Matt demonstrates high domain awareness by bringing up Alibaba's newly open-sourced Qwen 3 TTS model family. Neil explains how the industry builds on Moshi's open architecture and details Gradium's Pocket TTS running strictly on CPU.39:09–46:50 · Guest disagreement 3/10 Rethinking Voice Evaluation Beyond Traditional Benchmarks Matt pushes Neil on whether subjective human evaluation implies that voice models are becoming commoditized. Neil strongly rejects this premise, pointing out critical flaws in standard benchmarks like LibriSpeech and highlighting last-mile edge cases.46:50–49:39 · Guest disagreement 1/10 Cascaded Architecture vs. Direct Speech-to-Speech and Paralinguistics Neil explains the technical differences between cascaded architecture (ASR -> LLM -> TTS) and direct speech-to-speech. He details how text bottlenecks discard rich paralinguistic cues like emotion, sarcasm, and tone.49:39–53:29 · Guest disagreement 4/10 The Problem with Turn-Taking and Rule-Based Voice Activity Detection Neil forcefully critiques current turn-taking mechanisms, stating he hates rule-based Voice Activity Detection from the bottom of his heart. He explains how Moshi solved turn-taking using multi-stream token modeling on stereo datasets.53:29–56:14 · Guest disagreement 3/10 The Hardest Challenge in Voice AI: Noisy and Multi-Speaker Environments Neil makes a bold claim, betting that no speech team in the world will solve multi-speaker far-field voice recognition in noisy environments over the next year. He explains the psychoacoustics of binaural sound localization and unconscious human lip reading.56:14–1:00:50 · Guest disagreement 1/10 Voice AI Training Data: Quality, Diversity, and Script Generation Neil details training data dynamics, explaining why attempting to learn world knowledge from speech audio tokens rather than text is a terrible idea. He shares Gradium's intricate data engineering methods for generating random script prompts without model collapse.1:00:50–1:02:55 · Guest disagreement 1/10 Multilingual Transfer, Low-Resource Languages, and Parallel Datasets Neil outlines multilingual transfer learning across language families, noting that 1,000 hours was sufficient for Italian after training on Spanish and Portuguese. He points out that historical datasets like parallel Bible translations remain vital for low-resource languages.1:02:55–1:06:30 · Guest disagreement 1/10 Compute Economics, Hardware Constraints, and On-Demand Compute Scaling Neil details compute economics for voice agents, such as gaming NPCs in $70 video games where massive API costs are non-viable. He highlights Gradium's lean 15-person structure working alongside forward deployment partners.1:06:30–1:09:02 · Guest disagreement 1/10 Voice Cloning Capabilities vs. Natural Language Voice Design Neil differentiates voice cloning of specific IP/personalities from natural language voice design. He explains why enterprise applications lean heavily toward promptable voice design over exact audio replication.1:09:02–1:12:30 · Guest disagreement 5/10 Discussing Audio Watermarking, Deepfakes, and Voice Cloning Privacy Matt raises deepfake and privacy concerns surrounding voice cloning. Neil aggressively dismisses industry watermarking solutions as a scam, citing empirical breaking methods published in the Moshi paper.1:12:30–1:16:29 · Guest disagreement 1/10 Exploring Multimodal AI, Audio-Visual Integration, and Bridger Clone Matt prompts a discussion on multimodal integration between voice and video screens. Neil highlights audiovisual speech understanding and mentions consumer experiments like the Bridger Clone app.1:16:29–1:22:28 · Guest disagreement 4/10 Building Gradium in Paris and the French AI Ecosystem Matt brings up US tech cynicism toward European AI, citing President Macron's PR misfire regarding AI funding. Neil responds with fierce regional pride, defending Parisian talent density while mocking performative Silicon Valley culture like coding at the gym.0:00–3:27 · Matt pushing back 1/10 Episode Highlights: Voice AI's Turning Point Matt sets the stage regarding Voice AI's inflection point and asks Neil whether the field is truly at a major moment or still early. Neil collaboratively explains that while latency and accuracy have made huge strides, real-world deployment like noisy environments remains largely unsolved.3:27–7:39 · Matt pushing back 1/10 Historical Challenges and the Shortage of Voice AI Experts Matt asks why voice lagged behind text and vision historically. Neil educates Matt on the historical academic dynamics of speech research, noting that deep learning's first real success was actually speech recognition in 2007, despite speech lacking conference prestige.7:39–12:55 · Matt pushing back 4/10 Full-Duplex Architecture, Emotional Intelligence, and Future Interfaces Matt challenges the premise of ubiquitous voice AI by noting open office privacy constraints where people do not want to talk out loud. Neil reframes the problem, pointing out that voice coding and LLM prompting are shifting paradigms so fast that office layouts themselves may adapt.12:55–22:20 · Matt pushing back 1/10 Neil Zeghidour's Journey from Applied Math to Audio Language Models Neil walks through his background from applied math to Facebook AI, Google Brain, and neural audio codecs. Matt prompts for definitions, and Neil provides a masterclass on how neural codecs compress audio like text, enabling instant zero-shot voice cloning.22:20–27:09 · Matt pushing back 1/10 Leaving Big Tech and Founding Kyutai AI Lab Neil recounts founding Kyutai as an open-science non-profit lab and developing Moshi with a small team in six months. He explains why voice AI does not require 10,000 GPUs, giving small agile teams a distinct advantage over big tech.27:09–31:26 · Matt pushing back 1/10 Spinning Out Gradium from Open-Source Research Success Neil discusses spinning out Gradium for enterprise-grade product needs. He playfully describes his opportunistic research strategy of publishing open papers to frustrate closed-source rivals into joining his team.31:26–34:01 · Matt pushing back 3/10 Why Focused Startups Can Outmaneuver Big Tech in Voice Matt directly asks why big tech giants like OpenAI or Google have not already won voice AI. Neil breaks down parameter efficiency trade-offs in multimodal models and explains why focused developer building blocks beat single generalist assistants.34:01–39:09 · Matt pushing back 2/10 On-Device Voice AI and the Launch of Pocket TTS Matt demonstrates high domain awareness by bringing up Alibaba's newly open-sourced Qwen 3 TTS model family. Neil explains how the industry builds on Moshi's open architecture and details Gradium's Pocket TTS running strictly on CPU.39:09–46:50 · Matt pushing back 5/10 Rethinking Voice Evaluation Beyond Traditional Benchmarks Matt pushes Neil on whether subjective human evaluation implies that voice models are becoming commoditized. Neil strongly rejects this premise, pointing out critical flaws in standard benchmarks like LibriSpeech and highlighting last-mile edge cases.46:50–49:39 · Matt pushing back 1/10 Cascaded Architecture vs. Direct Speech-to-Speech and Paralinguistics Neil explains the technical differences between cascaded architecture (ASR -> LLM -> TTS) and direct speech-to-speech. He details how text bottlenecks discard rich paralinguistic cues like emotion, sarcasm, and tone.49:39–53:29 · Matt pushing back 1/10 The Problem with Turn-Taking and Rule-Based Voice Activity Detection Neil forcefully critiques current turn-taking mechanisms, stating he hates rule-based Voice Activity Detection from the bottom of his heart. He explains how Moshi solved turn-taking using multi-stream token modeling on stereo datasets.53:29–56:14 · Matt pushing back 1/10 The Hardest Challenge in Voice AI: Noisy and Multi-Speaker Environments Neil makes a bold claim, betting that no speech team in the world will solve multi-speaker far-field voice recognition in noisy environments over the next year. He explains the psychoacoustics of binaural sound localization and unconscious human lip reading.56:14–1:00:50 · Matt pushing back 1/10 Voice AI Training Data: Quality, Diversity, and Script Generation Neil details training data dynamics, explaining why attempting to learn world knowledge from speech audio tokens rather than text is a terrible idea. He shares Gradium's intricate data engineering methods for generating random script prompts without model collapse.1:00:50–1:02:55 · Matt pushing back 1/10 Multilingual Transfer, Low-Resource Languages, and Parallel Datasets Neil outlines multilingual transfer learning across language families, noting that 1,000 hours was sufficient for Italian after training on Spanish and Portuguese. He points out that historical datasets like parallel Bible translations remain vital for low-resource languages.1:02:55–1:06:30 · Matt pushing back 1/10 Compute Economics, Hardware Constraints, and On-Demand Compute Scaling Neil details compute economics for voice agents, such as gaming NPCs in $70 video games where massive API costs are non-viable. He highlights Gradium's lean 15-person structure working alongside forward deployment partners.1:06:30–1:09:02 · Matt pushing back 1/10 Voice Cloning Capabilities vs. Natural Language Voice Design Neil differentiates voice cloning of specific IP/personalities from natural language voice design. He explains why enterprise applications lean heavily toward promptable voice design over exact audio replication.1:09:02–1:12:30 · Matt pushing back 2/10 Discussing Audio Watermarking, Deepfakes, and Voice Cloning Privacy Matt raises deepfake and privacy concerns surrounding voice cloning. Neil aggressively dismisses industry watermarking solutions as a scam, citing empirical breaking methods published in the Moshi paper.1:12:30–1:16:29 · Matt pushing back 1/10 Exploring Multimodal AI, Audio-Visual Integration, and Bridger Clone Matt prompts a discussion on multimodal integration between voice and video screens. Neil highlights audiovisual speech understanding and mentions consumer experiments like the Bridger Clone app.1:16:29–1:22:28 · Matt pushing back 4/10 Building Gradium in Paris and the French AI Ecosystem Matt brings up US tech cynicism toward European AI, citing President Macron's PR misfire regarding AI funding. Neil responds with fierce regional pride, defending Parisian talent density while mocking performative Silicon Valley culture like coding at the gym.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 39.9% · guest 60.1%0:00 · Matt 39.9% · guest 60.1%3:00 · Matt 14.7% · guest 85.3%3:00 · Matt 14.7% · guest 85.3%6:00 · Matt 14.8% · guest 85.2%6:00 · Matt 14.8% · guest 85.2%9:00 · Matt 12.9% · guest 87.1%9:00 · Matt 12.9% · guest 87.1%12:00 · Matt 9.2% · guest 90.8%12:00 · Matt 9.2% · guest 90.8%15:00 · Matt 1.9% · guest 98.1%15:00 · Matt 1.9% · guest 98.1%18:00 · Matt 1.1% · guest 98.9%18:00 · Matt 1.1% · guest 98.9%21:00 · Matt 5.4% · guest 94.6%21:00 · Matt 5.4% · guest 94.6%24:00 · Matt 0% · guest 100%24:00 · Matt 0% · guest 100%27:00 · Matt 1.8% · guest 98.2%27:00 · Matt 1.8% · guest 98.2%30:00 · Matt 17.2% · guest 82.8%30:00 · Matt 17.2% · guest 82.8%33:00 · Matt 12.3% · guest 87.7%33:00 · Matt 12.3% · guest 87.7%36:00 · Matt 3.7% · guest 96.3%36:00 · Matt 3.7% · guest 96.3%39:00 · Matt 22.1% · guest 77.9%39:00 · Matt 22.1% · guest 77.9%42:00 · Matt 9.6% · guest 90.4%42:00 · Matt 9.6% · guest 90.4%45:00 · Matt 19.5% · guest 80.5%45:00 · Matt 19.5% · guest 80.5%48:00 · Matt 0% · guest 100%48:00 · Matt 0% · guest 100%51:00 · Matt 1% · guest 99%51:00 · Matt 1% · guest 99%54:00 · Matt 19.5% · guest 80.5%54:00 · Matt 19.5% · guest 80.5%57:00 · Matt 0% · guest 100%57:00 · Matt 0% · guest 100%1:00:00 · Matt 15.1% · guest 84.9%1:00:00 · Matt 15.1% · guest 84.9%1:03:00 · Matt 15.5% · guest 84.5%1:03:00 · Matt 15.5% · guest 84.5%1:06:00 · Matt 3.5% · guest 96.5%1:06:00 · Matt 3.5% · guest 96.5%1:09:00 · Matt 17.6% · guest 82.4%1:09:00 · Matt 17.6% · guest 82.4%1:12:00 · Matt 23.8% · guest 76.2%1:12:00 · Matt 23.8% · guest 76.2%1:15:00 · Matt 38% · guest 62%1:15:00 · Matt 38% · guest 62%1:18:00 · Matt 0% · guest 100%1:18:00 · Matt 0% · guest 100%1:21:00 · Matt 26.2% · guest 73.8%1:21:00 · Matt 26.2% · guest 73.8%
Sharpest disagreement ▶ 1:09:15 Calling audio watermarking a scam

Neil forcefully dismisses industry standard safety claims, calling audio watermarking an outright scam based on empirical breaks demonstrated in his research.

Hardest push from Matt ▶ 44:40 Challenging voice AI commoditization

Matt directly challenges the guest on whether the subjectivity of voice quality implies that the entire model layer is commoditized commodity tech.

Biggest teaching moment ▶ 57:00 World knowledge density in text versus speech

Neil educates the host on why pretraining speech models directly on audio for intelligence is fundamentally flawed compared to text-first architecture.

Matt holds his own ▶ 36:05 Citing Alibaba's Qwen 3 TTS release

Matt demonstrates deep expertise by citing Alibaba's freshly released Qwen 3 TTS open-source model family to test Gradium's competitive positioning.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Episode Highlights: Voice AI's Turning Point 3211 Matt sets the stage regarding Voice AI's inflection point and asks Neil whether the field is truly at a major moment or still early. Neil collaboratively explains that while latency and accuracy have made huge strides, real-world deployment like noisy environments remains largely unsolved.
Historical Challenges and the Shortage of Voice AI Experts 3511 Matt asks why voice lagged behind text and vision historically. Neil educates Matt on the historical academic dynamics of speech research, noting that deep learning's first real success was actually speech recognition in 2007, despite speech lacking conference prestige.
Full-Duplex Architecture, Emotional Intelligence, and Future Interfaces 4524 Matt challenges the premise of ubiquitous voice AI by noting open office privacy constraints where people do not want to talk out loud. Neil reframes the problem, pointing out that voice coding and LLM prompting are shifting paradigms so fast that office layouts themselves may adapt.
Neil Zeghidour's Journey from Applied Math to Audio Language Models 3611 Neil walks through his background from applied math to Facebook AI, Google Brain, and neural audio codecs. Matt prompts for definitions, and Neil provides a masterclass on how neural codecs compress audio like text, enabling instant zero-shot voice cloning.
Leaving Big Tech and Founding Kyutai AI Lab 3411 Neil recounts founding Kyutai as an open-science non-profit lab and developing Moshi with a small team in six months. He explains why voice AI does not require 10,000 GPUs, giving small agile teams a distinct advantage over big tech.
Spinning Out Gradium from Open-Source Research Success 3421 Neil discusses spinning out Gradium for enterprise-grade product needs. He playfully describes his opportunistic research strategy of publishing open papers to frustrate closed-source rivals into joining his team.
Why Focused Startups Can Outmaneuver Big Tech in Voice 4523 Matt directly asks why big tech giants like OpenAI or Google have not already won voice AI. Neil breaks down parameter efficiency trade-offs in multimodal models and explains why focused developer building blocks beat single generalist assistants.
On-Device Voice AI and the Launch of Pocket TTS 4522 Matt demonstrates high domain awareness by bringing up Alibaba's newly open-sourced Qwen 3 TTS model family. Neil explains how the industry builds on Moshi's open architecture and details Gradium's Pocket TTS running strictly on CPU.
Rethinking Voice Evaluation Beyond Traditional Benchmarks 5635 Matt pushes Neil on whether subjective human evaluation implies that voice models are becoming commoditized. Neil strongly rejects this premise, pointing out critical flaws in standard benchmarks like LibriSpeech and highlighting last-mile edge cases.
Cascaded Architecture vs. Direct Speech-to-Speech and Paralinguistics 3611 Neil explains the technical differences between cascaded architecture (ASR -> LLM -> TTS) and direct speech-to-speech. He details how text bottlenecks discard rich paralinguistic cues like emotion, sarcasm, and tone.
The Problem with Turn-Taking and Rule-Based Voice Activity Detection 3641 Neil forcefully critiques current turn-taking mechanisms, stating he hates rule-based Voice Activity Detection from the bottom of his heart. He explains how Moshi solved turn-taking using multi-stream token modeling on stereo datasets.
The Hardest Challenge in Voice AI: Noisy and Multi-Speaker Environments 3631 Neil makes a bold claim, betting that no speech team in the world will solve multi-speaker far-field voice recognition in noisy environments over the next year. He explains the psychoacoustics of binaural sound localization and unconscious human lip reading.
Voice AI Training Data: Quality, Diversity, and Script Generation 3611 Neil details training data dynamics, explaining why attempting to learn world knowledge from speech audio tokens rather than text is a terrible idea. He shares Gradium's intricate data engineering methods for generating random script prompts without model collapse.
Multilingual Transfer, Low-Resource Languages, and Parallel Datasets 3511 Neil outlines multilingual transfer learning across language families, noting that 1,000 hours was sufficient for Italian after training on Spanish and Portuguese. He points out that historical datasets like parallel Bible translations remain vital for low-resource languages.
Compute Economics, Hardware Constraints, and On-Demand Compute Scaling 3511 Neil details compute economics for voice agents, such as gaming NPCs in $70 video games where massive API costs are non-viable. He highlights Gradium's lean 15-person structure working alongside forward deployment partners.
Voice Cloning Capabilities vs. Natural Language Voice Design 3511 Neil differentiates voice cloning of specific IP/personalities from natural language voice design. He explains why enterprise applications lean heavily toward promptable voice design over exact audio replication.
Discussing Audio Watermarking, Deepfakes, and Voice Cloning Privacy 4652 Matt raises deepfake and privacy concerns surrounding voice cloning. Neil aggressively dismisses industry watermarking solutions as a scam, citing empirical breaking methods published in the Moshi paper.
Exploring Multimodal AI, Audio-Visual Integration, and Bridger Clone 4511 Matt prompts a discussion on multimodal integration between voice and video screens. Neil highlights audiovisual speech understanding and mentions consumer experiments like the Bridger Clone app.
Building Gradium in Paris and the French AI Ecosystem 5544 Matt brings up US tech cynicism toward European AI, citing President Macron's PR misfire regarding AI funding. Neil responds with fierce regional pride, defending Parisian talent density while mocking performative Silicon Valley culture like coding at the gym.

Statements from this episode (38)

Opinion
Zeghidour: Phone calls with AI agents can now outperform human interactions
“For the first time it's actually can be enjoyable to, and even more convenient to talk to an AI on the phones and talking to a human because you can call any time of the day or night and the interaction is is working pretty well and it sounds really nice and t…”
Neil Zeghidour Feb 19, 2026 ▶ 2:22
Assertion Supported
Zeghidour: Deep learning's first major success was speech recognition, preceding AlexNet
“Actually the really first big success in deep learning was speech recognition. We've worked from Geoffrey Hinton himself.”
Neil Zeghidour Feb 19, 2026 ▶ 4:52
Assertion Not checkable as stated
Zeghidour: Only 50 people worldwide can train competitive voice AI models
“Between 10 and 100? No, I would say. 50? I don't know. It's hard to say. But, yeah, I think it's very few and, really meaningful contributions that have pushed the field forward have been made by very small groups of people.”
Neil Zeghidour Feb 19, 2026 ▶ 6:55
Insight
Zeghidour: Voice AI's lower compute and data requirements enable small teams
“And in voice in particular, since the required compute is much lower and is that the same for data, really a few individuals can make stuff that is completely You know, just changing applications at very large scales.”
Neil Zeghidour Feb 19, 2026 ▶ 7:26
Prediction Not checkable as stated
Zeghidour: Voice will be the primary interface for next-gen AI hardware
“In my perception, all the new hardware companies have voice at the heart of the product. All the prototypes that we see, whether it's glasses or pendants or, you know, like the new stuff that Johnny Hive and Sam Altman are working on. Voice is at the heart of …”
Neil Zeghidour Feb 19, 2026 ▶ 10:30
Assertion Not checkable as stated
Zeghidour: Speech AI was largely considered a solved problem at Google Brain in 2019
“At that time, it was interesting because so I joined working on speech in Google Brain, and there were almost nobody working on speech in Google Brain. It was not considered vibrant research topic. It was like a product topic. A lot of people were saying, oh, …”
Neil Zeghidour Feb 19, 2026 ▶ 16:16
Assertion Not checkable as stated
Zeghidour: Audio language models dominate voice AI due to streaming capabilities
“I think today, virtually everything is audio language models because since they are autoregressive, so they run in in a streaming fashion, they are naturally compatible with real-time inference, which is kind of the main topic around voice right now. And so ev…”
Neil Zeghidour Feb 19, 2026 ▶ 21:22
Assertion Contradicted
Zeghidour: Kyutai's Moshi remains the only full-duplex conversational AI model
“Moshi, that is still to the day the only full duplex model.”
Neil Zeghidour Feb 19, 2026 ▶ 25:45
Opinion
Zeghidour: No outside team was capable of commercializing Kyutai's voice models
“And honestly, after a few interactions, I realized nobody could carry such a project except us.”
Neil Zeghidour Feb 19, 2026 ▶ 28:12
Opinion
Zeghidour: Chinese open research forces Western competitors to publish
“And they also, you know, like the Chinese lab, I, I'm making a remarkable work, and it's a kind of, are forcing everyone to stay open to some extent because otherwise it also hurts the ego, I think, of the people who are in the labs that don't publish.”
Neil Zeghidour Feb 19, 2026 ▶ 30:28
Assertion Not checkable as stated
Zeghidour: Large multimodal models are too massive to run voice profitably
“And at the same time, these models are so large, they cannot run at scale because they will just make everyone lose money in the process.”
Neil Zeghidour Feb 19, 2026 ▶ 32:37
Opinion
Zeghidour: Big Tech lacks the DNA to build targeted developer-focused models
“And this again is, it's not really, I think, in the DNA of big companies to do this kind of very specific models that are targeted towards developers, rather than trying to solve a lot of things at the same time.”
Neil Zeghidour Feb 19, 2026 ▶ 33:49
Insight
On-Device Voice AI Enables Large-Scale Personalization Unfeasible via Cloud APIs
“On device models allow to do very large scale personalized content that will be economically not realistic with an API.”
Neil Zeghidour Feb 19, 2026 ▶ 35:07
Insight
Zeghidour: Building Small Voice AI Models Is Harder Than Large Ones
“Making small models in voice is much more difficult than making large models in voice. So keeping the quality while reducing the size of the model, that's where the big challenge is.”
Neil Zeghidour Feb 19, 2026 ▶ 35:23
Assertion Supported
Zeghidour: Mistral and Alibaba's voice models build upon Kyutai's Moshi architecture
“It's mostly inspired from the Moshi architecture, like pretty much every model right now, even the Voxstral model that was released by Mistral two weeks ago is also based on on our framework.”
Neil Zeghidour Feb 19, 2026 ▶ 36:12
Assertion Not checkable as stated
Zeghidour: Neural network proxies for audio grading fail on real-world audio
“People have tried to make objective proxies of human judgment. Like that would be a neural network that listens to an audio and gives it a grade. It sucks. Like so many people try and it works on their constrained setting and on real audio it doesn't work at a…”
Neil Zeghidour Feb 19, 2026 ▶ 42:02
Assertion Not checkable as stated
Zeghidour: Speaker diarization models completely break on multi-speaker podcasts
“At the same time, it's a extremely useful problem. And you look at the error rates and they are very bad. I mean, it's still just not working in difficult cases where you have a podcast with a lot of people talking at the same time, just completely breaks.”
Neil Zeghidour Feb 19, 2026 ▶ 45:43
Prediction Not checkable as stated
Zeghidour: Voice AI is very far from becoming commoditized
“Full duplex. We, you know, we did Moshi a year and a half ago. Still nobody has made it into a product. There are so many things that are just not existing today that I think the communitization, maybe it will happen someday, but we are very, very far from it,…”
Neil Zeghidour Feb 19, 2026 ▶ 46:00
Insight
Zeghidour: Cascaded voice AI loses paralinguistic info in text bottleneck
“By going through the bottleneck of text, you lose what we call paralinguistic information, which is all the information we convey when we speak on top of what we say.”
Neil Zeghidour Feb 19, 2026 ▶ 48:08
Opinion
Zeghidour: Voice AI turn-taking relies on archaic, handmade rules
“With turn taking, we are back at the archaic era of handmade rules, which is ridiculous.”
Neil Zeghidour Feb 19, 2026 ▶ 50:10
Opinion
Zeghidour: Users must adapt to the flow of current cascaded voice AI
“You need discipline when you talk to AI. You need to adapt to its flow. Otherwise it's, it gets lost and gets confused and interrupts and so on.”
Neil Zeghidour Feb 19, 2026 ▶ 50:39
Insight
Zeghidour: End-to-end speech models have massive switching costs for upgrades
“One drawback of speech-to-speech models is that since everything is integrated, when you go from a text model to the speech-to-speech model, you need to fine-tune it on speech data. So now it's the cost to switch the underlying text model is extremely high bec…”
Neil Zeghidour Feb 19, 2026 ▶ 52:50
Prediction Not checkable as stated
Zeghidour: No AI team will solve noisy multi-speaker recognition within a year
“Well, I would say a frontier is I could like bet to every single speech team in the world that they don't solve it in the next year or so. It's a robot in the model in the factory, and there is a lot of noise from machines, and you have a lot of people talking…”
Neil Zeghidour Feb 19, 2026 ▶ 53:34
Assertion Not checkable as stated
Zeghidour: Zero progress made on noisy multi-speaker understanding in 10 years
“In TTS, we see a lot of progress. For this kind of hard understanding problem, I hear people saying the exact same stuff as they did 10 years ago. Like, there was zero progress.”
Neil Zeghidour Feb 19, 2026 ▶ 56:03
Insight
Zeghidour: Training AI to learn world knowledge from speech is terrible
“Getting your model to learn about the world from speech, I think it's a terrible idea.”
Neil Zeghidour Feb 19, 2026 ▶ 57:17
Disclosure
Zeghidour: Kyutai trained its Moshi model on 7 million hours of speech
“So for Moshi, we trained on seven million hours of speech.”
Neil Zeghidour Feb 19, 2026 ▶ 57:48
Insight
Zeghidour: LLMs cannot generate massive diverse synthetic voice scripts without collapsing
“I mean, you cannot ask Claude or ChatGPT write 100,000 hours of scripts and make them as diverse as possible. So it doesn't work. It's going just to be In a loop and collapse on a few topics, you know.”
Neil Zeghidour Feb 19, 2026 ▶ 59:06
Assertion Supported
Hibiki Zero Needed Only 1,000 Italian Hours After Spanish Pretraining
“For example, we have Hibiki Zero that released last week. So it was trained on 50,000 hours or maybe 100,000 hours for example, Portuguese and Spanish. And then to do Italian translation to English, 1000 hours was enough.”
Neil Zeghidour Feb 19, 2026 ▶ 1:01:18
Insight
Zeghidour: Selective compute usage is essential for voice AI economics
“So being very selective about when to compute, to use compute, I think that's the only way to, for all of it to make sense economically.”
Neil Zeghidour Feb 19, 2026 ▶ 1:04:03
Opinion
Zeghidour: Voice design makes more sense than cloning for corporate voice AI
“For these customers, I think the solution that makes the most sense is not cloning. It's voice design.”
Neil Zeghidour Feb 19, 2026 ▶ 1:08:04
Assertion Not checkable as stated
Zeghidour: Existing voice design tools fail because they lack precise control
“There are a few solutions that exist today around voice design, but as far as I know, they are not very popular and people still stick to the existing voice catalog because they cannot steer it precisely enough.”
Neil Zeghidour Feb 19, 2026 ▶ 1:08:17
Opinion
Zeghidour: Audio watermarking is a scam and easily broken
“Watermarking is a scam. I'm sorry, I have to say it. It just doesn't work. I worked on it. We have an appendix in the Moshi paper around how we could break so easily any watermarking that was supposed to be state of the art. So people should not rely on that.”
Neil Zeghidour Feb 19, 2026 ▶ 1:09:16
Prediction Not checkable as stated
Zeghidour: AI voice design will eliminate the need for voice cloning
“Voice design is going to, you know, just remove this issue because then again, people typically are going to clone the voice of someone, but what they wanted is someone from a specific gender, specific demographics, age, accent, and so on. And so they could ju…”
Neil Zeghidour Feb 19, 2026 ▶ 1:12:08
Opinion
Zeghidour: Generating video and audio separately ignores inherent multimodal training data
“I don't think it really makes sense to do video to audio generation separately, because typically the data exists as a multimodal signal, right?”
Neil Zeghidour Feb 19, 2026 ▶ 1:13:29
Opinion
Zeghidour: European AI is mostly centered in France
“European AI is mostly French AI, to be fair. There is also Germany, but a lot of it is in France.”
Neil Zeghidour Feb 19, 2026 ▶ 1:18:27
Assertion Supported
Zeghidour: Most of Google's audio AI models were built in Paris and Zurich
“A lot of the current audio generative models of Google, a global company, obviously were developed between Paris and Zurich. Actually most of it.”
Neil Zeghidour Feb 19, 2026 ▶ 1:18:40
Assertion Supported
Zeghidour: Meta's LLaMA and DINO models were developed in Paris
“Lama was started in Paris. Dino, which is the most groundbreaking vision work from Facebook was developed in Paris.”
Neil Zeghidour Feb 19, 2026 ▶ 1:18:50
Opinion
Zeghidour: Mistral's top researchers match the best talent at major US labs
“I know The best people from Mistral, they are, you know, they can be compared to the top of the top of the biggest labs.”
Neil Zeghidour Feb 19, 2026 ▶ 1:21:36
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.