Text To Speech
topic on 7 shows · 12 statements across 10 episodes
the Y Combinator Startup Podcast
David Senra
Latent Space
the Neon Show
No Priors
the MAD Podcast
TBPN
12 statements about Text To Speech, every show
Staniszewski claims ElevenLabs built the first human-quality speech model in 2022
“Built all the audio models, starting with model to produce speech text to speech model. And that was that 20, 20 to the first model that could finally cross that human like quality.”
Staniszewski: Benchmarking text-to-speech models is extremely difficult due to voice differences
“So like even doing benchmarks for text to speech is extremely hard. Because usually different models will have different voices. That already makes them uncomparable.”
Kamath: Smallest AI cut TTS latency to sub-100ms by seed round
“So we had a text to speech model which basically you could use to give natural realistic voice to the voice agent in under a hundred milliseconds. So initially it was at 200 milliseconds. It came down to a hundred milliseconds by the time our seed round was th…”
Zeghidour: Zero progress made on noisy multi-speaker understanding in 10 years
“In TTS, we see a lot of progress. For this kind of hard understanding problem, I hear people saying the exact same stuff as they did 10 years ago. Like, there was zero progress.”
Hsu: Only a few TTS models handle multilingual code-switching properly
“It's actually like only a few models are able to speak two languages in the same sentence and then pronounce them properly.”
Russ d'Sa: LLM inference is now faster than text-to-speech generation
“Now like LLM inference can actually be done in less time than generating speech with TTS.”
D'Sa: LLM inference is now faster than TTS speech generation
“And now like LLM inference can actually be done in less time than generating speech with TTS.”
Martin: Dual Personas With Editorial Takes Make AI Audio Engaging
“There is a transform that needs to happen. That is inherently editorial. And I think this is where like that two person persona, right? Dialogue model, they have takes on the material that you've presented. That's where it really sort of like brings the conten…”
Goel: Current AI speech models fail to capture profession-specific vocal nuances
“And that's sort of the nuance that I don't think any models really capture well, which is like, you know, if you're a nurse, you need to talk in a different way than if you're a lawyer, or if you're a judge, or if you're a venture capitalist, you know, very di…”
Gu: High-quality speech synthesis requires multimodal foundation models
“And so actually to really get, like, perfect even just TTS or, like, speech-to-speech you actually really need to have, like, a model that has, More understanding, like at least of the language, but kind of like, it's not really an isolated component anymore. …”
Shulman: Bark was the first open-source transformer-based TTS model
“As far as I know there was no other certainly not in the open source text to speech that was kind of transformer based.”
Pre-training on thousands of voices enables low-data voice cloning
“If I learn to mimic lots of different voices and then you give me the 1001st voice you'd hope that the first thousand taught you virtually everything you need to know about language and that what's left is really some idiosyncratic change. That you could learn…”