The Ledger, every show
Every statement that passed quotation and attribution checks, across all 44 shows. Pick shows below, then mix any filter with any other.
shows 




every show 44 of 44
Zeghidour: Chinese open research forces Western competitors to publish
“And they also, you know, like the Chinese lab, I, I'm making a remarkable work, and it's a kind of, are forcing everyone to stay open to some extent because otherwise it also hurts the ego, I think, of the people who are in the labs that don't publish.”
Zeghidour: Zero progress made on noisy multi-speaker understanding in 10 years
“In TTS, we see a lot of progress. For this kind of hard understanding problem, I hear people saying the exact same stuff as they did 10 years ago. Like, there was zero progress.”
Zeghidour: Audio watermarking is a scam and easily broken
“Watermarking is a scam. I'm sorry, I have to say it. It just doesn't work. I worked on it. We have an appendix in the Moshi paper around how we could break so easily any watermarking that was supposed to be state of the art. So people should not rely on that.”
Zeghidour: Only 50 people worldwide can train competitive voice AI models
“Between 10 and 100? No, I would say. 50? I don't know. It's hard to say. But, yeah, I think it's very few and, really meaningful contributions that have pushed the field forward have been made by very small groups of people.”
Zeghidour: Voice will be the primary interface for next-gen AI hardware
“In my perception, all the new hardware companies have voice at the heart of the product. All the prototypes that we see, whether it's glasses or pendants or, you know, like the new stuff that Johnny Hive and Sam Altman are working on. Voice is at the heart of …”
Zeghidour: Kyutai's Moshi remains the only full-duplex conversational AI model
“Moshi, that is still to the day the only full duplex model.”
Zeghidour: Large multimodal models are too massive to run voice profitably
“And at the same time, these models are so large, they cannot run at scale because they will just make everyone lose money in the process.”
Zeghidour: Voice AI is very far from becoming commoditized
“Full duplex. We, you know, we did Moshi a year and a half ago. Still nobody has made it into a product. There are so many things that are just not existing today that I think the communitization, maybe it will happen someday, but we are very, very far from it,…”
Zeghidour: No AI team will solve noisy multi-speaker recognition within a year
“Well, I would say a frontier is I could like bet to every single speech team in the world that they don't solve it in the next year or so. It's a robot in the model in the factory, and there is a lot of noise from machines, and you have a lot of people talking…”
Zeghidour: Training AI to learn world knowledge from speech is terrible
“Getting your model to learn about the world from speech, I think it's a terrible idea.”
Zeghidour: AI voice design will eliminate the need for voice cloning
“Voice design is going to, you know, just remove this issue because then again, people typically are going to clone the voice of someone, but what they wanted is someone from a specific gender, specific demographics, age, accent, and so on. And so they could ju…”
Zeghidour: Phone calls with AI agents can now outperform human interactions
“For the first time it's actually can be enjoyable to, and even more convenient to talk to an AI on the phones and talking to a human because you can call any time of the day or night and the interaction is is working pretty well and it sounds really nice and t…”
Zeghidour: Audio language models dominate voice AI due to streaming capabilities
“I think today, virtually everything is audio language models because since they are autoregressive, so they run in in a streaming fashion, they are naturally compatible with real-time inference, which is kind of the main topic around voice right now. And so ev…”
Zeghidour: No outside team was capable of commercializing Kyutai's voice models
“And honestly, after a few interactions, I realized nobody could carry such a project except us.”
Zeghidour: Big Tech lacks the DNA to build targeted developer-focused models
“And this again is, it's not really, I think, in the DNA of big companies to do this kind of very specific models that are targeted towards developers, rather than trying to solve a lot of things at the same time.”
Zeghidour: Building Small Voice AI Models Is Harder Than Large Ones
“Making small models in voice is much more difficult than making large models in voice. So keeping the quality while reducing the size of the model, that's where the big challenge is.”
Zeghidour: Neural network proxies for audio grading fail on real-world audio
“People have tried to make objective proxies of human judgment. Like that would be a neural network that listens to an audio and gives it a grade. It sucks. Like so many people try and it works on their constrained setting and on real audio it doesn't work at a…”
Zeghidour: Voice AI turn-taking relies on archaic, handmade rules
“With turn taking, we are back at the archaic era of handmade rules, which is ridiculous.”
Zeghidour: End-to-end speech models have massive switching costs for upgrades
“One drawback of speech-to-speech models is that since everything is integrated, when you go from a text model to the speech-to-speech model, you need to fine-tune it on speech data. So now it's the cost to switch the underlying text model is extremely high bec…”
Zeghidour: Voice design makes more sense than cloning for corporate voice AI
“For these customers, I think the solution that makes the most sense is not cloning. It's voice design.”
Zeghidour: Generating video and audio separately ignores inherent multimodal training data
“I don't think it really makes sense to do video to audio generation separately, because typically the data exists as a multimodal signal, right?”
Zeghidour: European AI is mostly centered in France
“European AI is mostly French AI, to be fair. There is also Germany, but a lot of it is in France.”
Zeghidour: Deep learning's first major success was speech recognition, preceding AlexNet
“Actually the really first big success in deep learning was speech recognition. We've worked from Geoffrey Hinton himself.”
Zeghidour: Voice AI's lower compute and data requirements enable small teams
“And in voice in particular, since the required compute is much lower and is that the same for data, really a few individuals can make stuff that is completely You know, just changing applications at very large scales.”
Zeghidour: Speech AI was largely considered a solved problem at Google Brain in 2019
“At that time, it was interesting because so I joined working on speech in Google Brain, and there were almost nobody working on speech in Google Brain. It was not considered vibrant research topic. It was like a product topic. A lot of people were saying, oh, …”
On-Device Voice AI Enables Large-Scale Personalization Unfeasible via Cloud APIs
“On device models allow to do very large scale personalized content that will be economically not realistic with an API.”
Zeghidour: Mistral and Alibaba's voice models build upon Kyutai's Moshi architecture
“It's mostly inspired from the Moshi architecture, like pretty much every model right now, even the Voxstral model that was released by Mistral two weeks ago is also based on on our framework.”
Zeghidour: Speaker diarization models completely break on multi-speaker podcasts
“At the same time, it's a extremely useful problem. And you look at the error rates and they are very bad. I mean, it's still just not working in difficult cases where you have a podcast with a lot of people talking at the same time, just completely breaks.”
Zeghidour: Users must adapt to the flow of current cascaded voice AI
“You need discipline when you talk to AI. You need to adapt to its flow. Otherwise it's, it gets lost and gets confused and interrupts and so on.”
Zeghidour: Kyutai trained its Moshi model on 7 million hours of speech
“So for Moshi, we trained on seven million hours of speech.”
Zeghidour: LLMs cannot generate massive diverse synthetic voice scripts without collapsing
“I mean, you cannot ask Claude or ChatGPT write 100,000 hours of scripts and make them as diverse as possible. So it doesn't work. It's going just to be In a loop and collapse on a few topics, you know.”
Zeghidour: Selective compute usage is essential for voice AI economics
“So being very selective about when to compute, to use compute, I think that's the only way to, for all of it to make sense economically.”
Zeghidour: Existing voice design tools fail because they lack precise control
“There are a few solutions that exist today around voice design, but as far as I know, they are not very popular and people still stick to the existing voice catalog because they cannot steer it precisely enough.”
Zeghidour: Mistral's top researchers match the best talent at major US labs
“I know The best people from Mistral, they are, you know, they can be compared to the top of the top of the biggest labs.”
Zeghidour: Cascaded voice AI loses paralinguistic info in text bottleneck
“By going through the bottleneck of text, you lose what we call paralinguistic information, which is all the information we convey when we speak on top of what we say.”
Hibiki Zero Needed Only 1,000 Italian Hours After Spanish Pretraining
“For example, we have Hibiki Zero that released last week. So it was trained on 50,000 hours or maybe 100,000 hours for example, Portuguese and Spanish. And then to do Italian translation to English, 1000 hours was enough.”