Text To Speech

topic on 7 shows · 12 statements across 10 episodes

the Y Combinator Startup Podcast David Senra Latent Space the Neon Show No Priors the MAD Podcast TBPN

12 statements about Text To Speech, every show

DAVID SENRA Assertion Not checkable as stated
Staniszewski claims ElevenLabs built the first human-quality speech model in 2022
“Built all the audio models, starting with model to produce speech text to speech model. And that was that 20, 20 to the first model that could finally cross that human like quality.”
Mati Staniszewski Sep 9, 2026 ▶ 2:22 Building One of AI’s Fastest-Growing Companies | Mati Staniszewski, ElevenLabs
Staniszewski: Benchmarking text-to-speech models is extremely difficult due to voice differences
“So like even doing benchmarks for text to speech is extremely hard. Because usually different models will have different voices. That already makes them uncomparable.”
Mati Staniszewski Sep 9, 2026 ▶ 55:09 Building One of AI’s Fastest-Growing Companies | Mati Staniszewski, ElevenLabs
NEON SHOW Assertion Not checkable as stated
Kamath: Smallest AI cut TTS latency to sub-100ms by seed round
“So we had a text to speech model which basically you could use to give natural realistic voice to the voice agent in under a hundred milliseconds. So initially it was at 200 milliseconds. It came down to a hundred milliseconds by the time our seed round was th…”
Sudarshan Kamath Mar 6, 2026 ▶ 21:21 Where SMALL models will Win | Sudarshan kamath, Smallest ai
MAD Assertion Not checkable as stated
Zeghidour: Zero progress made on noisy multi-speaker understanding in 10 years
“In TTS, we see a lot of progress. For this kind of hard understanding problem, I hear people saying the exact same stuff as they did 10 years ago. Like, there was zero progress.”
Neil Zeghidour Feb 19, 2026 ▶ 56:03 Voice AI’s Big Moment: Top Researcher on Why Everything Is Changing (Neil Zeghidour, Gradium AI)
LATENT SPACE Assertion Not checkable as stated
Hsu: Only a few TTS models handle multilingual code-switching properly
“It's actually like only a few models are able to speak two languages in the same sentence and then pronounce them properly.”
Andrew Hsu Jul 11, 2025 ▶ 47:33 Personalized AI Language Education — with Andrew Hsu, Speak
TBPN Assertion Supported
Russ d'Sa: LLM inference is now faster than text-to-speech generation
“Now like LLM inference can actually be done in less time than generating speech with TTS.”
Russell D'Sa Apr 26, 2025 ▶ 25:54 Building the World’s Smallest Fusion Reactor | Robin Langtry on TBPN
TBPN Assertion Supported
D'Sa: LLM inference is now faster than TTS speech generation
“And now like LLM inference can actually be done in less time than generating speech with TTS.”
Russell D'Sa Apr 26, 2025 ▶ 5:38 Why Hallucination Is Good For AI Voice Agents | Russell D'Sa on TBPN
Martin: Dual Personas With Editorial Takes Make AI Audio Engaging
“There is a transform that needs to happen. That is inherently editorial. And I think this is where like that two person persona, right? Dialogue model, they have takes on the material that you've presented. That's where it really sort of like brings the conten…”
Raiza Martin Oct 25, 2024 ▶ 15:43 How NotebookLM Was Made
NO PRIORS Opinion
Goel: Current AI speech models fail to capture profession-specific vocal nuances
“And that's sort of the nuance that I don't think any models really capture well, which is like, you know, if you're a nurse, you need to talk in a different way than if you're a lawyer, or if you're a judge, or if you're a venture capitalist, you know, very di…”
Karan Goel Jun 27, 2024 ▶ 23:42 No Priors Ep. 70 | With Cartesia Co-Founders Karan Goel & Albert Gu
NO PRIORS Insight
Gu: High-quality speech synthesis requires multimodal foundation models
“And so actually to really get, like, perfect even just TTS or, like, speech-to-speech you actually really need to have, like, a model that has, More understanding, like at least of the language, but kind of like, it's not really an isolated component anymore. …”
Albert Gu Jun 27, 2024 ▶ 24:27 No Priors Ep. 70 | With Cartesia Co-Founders Karan Goel & Albert Gu
LATENT SPACE Assertion Contradicted
Shulman: Bark was the first open-source transformer-based TTS model
“As far as I know there was no other certainly not in the open source text to speech that was kind of transformer based.”
Mikey Shulman Mar 14, 2024 ▶ 15:48 Making Transformers Sing - with Mikey Shulman of Suno
Pre-training on thousands of voices enables low-data voice cloning
“If I learn to mimic lots of different voices and then you give me the 1001st voice you'd hope that the first thousand taught you virtually everything you need to know about language and that what's left is really some idiosyncratic change. That you could learn…”
Adam Coates Aug 11, 2017 ▶ 6:59 Baidu's AI Lab Director on Advancing Speech Recognition and Simulation · Y Combinator

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.