Opinion
Zeghidour: Chinese open research forces Western competitors to publish
“And they also, you know, like the Chinese lab, I, I'm making a remarkable work, and it's a kind of, are forcing everyone to stay open to some extent because otherwise it also hurts the ego, I think, of the people who are in the labs that don't publish.”
Assertion Not checkable as stated
Zeghidour: Zero progress made on noisy multi-speaker understanding in 10 years
“In TTS, we see a lot of progress. For this kind of hard understanding problem, I hear people saying the exact same stuff as they did 10 years ago. Like, there was zero progress.”
Opinion
Zeghidour: Audio watermarking is a scam and easily broken
“Watermarking is a scam. I'm sorry, I have to say it. It just doesn't work. I worked on it. We have an appendix in the Moshi paper around how we could break so easily any watermarking that was supposed to be state of the art. So people should not rely on that.”
Assertion Not checkable as stated
Zeghidour: Only 50 people worldwide can train competitive voice AI models
“Between 10 and 100? No, I would say. 50? I don't know. It's hard to say. But, yeah, I think it's very few and, really meaningful contributions that have pushed the field forward have been made by very small groups of people.”
Prediction Not checkable as stated
Zeghidour: Voice will be the primary interface for next-gen AI hardware
“In my perception, all the new hardware companies have voice at the heart of the product. All the prototypes that we see, whether it's glasses or pendants or, you know, like the new stuff that Johnny Hive and Sam Altman are working on. Voice is at the heart of …”
Assertion Contradicted
Zeghidour: Kyutai's Moshi remains the only full-duplex conversational AI model
“Moshi, that is still to the day the only full duplex model.”
Assertion Not checkable as stated
Zeghidour: Large multimodal models are too massive to run voice profitably
“And at the same time, these models are so large, they cannot run at scale because they will just make everyone lose money in the process.”
Prediction Not checkable as stated
Zeghidour: Voice AI is very far from becoming commoditized
“Full duplex. We, you know, we did Moshi a year and a half ago. Still nobody has made it into a product. There are so many things that are just not existing today that I think the communitization, maybe it will happen someday, but we are very, very far from it,…”
Prediction Not checkable as stated
Zeghidour: No AI team will solve noisy multi-speaker recognition within a year
“Well, I would say a frontier is I could like bet to every single speech team in the world that they don't solve it in the next year or so. It's a robot in the model in the factory, and there is a lot of noise from machines, and you have a lot of people talking…”
Insight
Zeghidour: Training AI to learn world knowledge from speech is terrible
“Getting your model to learn about the world from speech, I think it's a terrible idea.”
Prediction Not checkable as stated
Zeghidour: AI voice design will eliminate the need for voice cloning
“Voice design is going to, you know, just remove this issue because then again, people typically are going to clone the voice of someone, but what they wanted is someone from a specific gender, specific demographics, age, accent, and so on. And so they could ju…”
Opinion
Zeghidour: Phone calls with AI agents can now outperform human interactions
“For the first time it's actually can be enjoyable to, and even more convenient to talk to an AI on the phones and talking to a human because you can call any time of the day or night and the interaction is is working pretty well and it sounds really nice and t…”
Assertion Not checkable as stated
Zeghidour: Audio language models dominate voice AI due to streaming capabilities
“I think today, virtually everything is audio language models because since they are autoregressive, so they run in in a streaming fashion, they are naturally compatible with real-time inference, which is kind of the main topic around voice right now. And so ev…”
Opinion
Zeghidour: No outside team was capable of commercializing Kyutai's voice models
“And honestly, after a few interactions, I realized nobody could carry such a project except us.”
Opinion
Zeghidour: Big Tech lacks the DNA to build targeted developer-focused models
“And this again is, it's not really, I think, in the DNA of big companies to do this kind of very specific models that are targeted towards developers, rather than trying to solve a lot of things at the same time.”
Insight
Zeghidour: Building Small Voice AI Models Is Harder Than Large Ones
“Making small models in voice is much more difficult than making large models in voice. So keeping the quality while reducing the size of the model, that's where the big challenge is.”
Assertion Not checkable as stated
Zeghidour: Neural network proxies for audio grading fail on real-world audio
“People have tried to make objective proxies of human judgment. Like that would be a neural network that listens to an audio and gives it a grade. It sucks. Like so many people try and it works on their constrained setting and on real audio it doesn't work at a…”
Opinion
Zeghidour: Voice AI turn-taking relies on archaic, handmade rules
“With turn taking, we are back at the archaic era of handmade rules, which is ridiculous.”
Insight
Zeghidour: End-to-end speech models have massive switching costs for upgrades
“One drawback of speech-to-speech models is that since everything is integrated, when you go from a text model to the speech-to-speech model, you need to fine-tune it on speech data. So now it's the cost to switch the underlying text model is extremely high bec…”
Opinion
Zeghidour: Voice design makes more sense than cloning for corporate voice AI
“For these customers, I think the solution that makes the most sense is not cloning. It's voice design.”
Opinion
Zeghidour: Generating video and audio separately ignores inherent multimodal training data
“I don't think it really makes sense to do video to audio generation separately, because typically the data exists as a multimodal signal, right?”
Opinion
Zeghidour: European AI is mostly centered in France
“European AI is mostly French AI, to be fair. There is also Germany, but a lot of it is in France.”
Assertion Supported
Zeghidour: Deep learning's first major success was speech recognition, preceding AlexNet
“Actually the really first big success in deep learning was speech recognition. We've worked from Geoffrey Hinton himself.”
Insight
Zeghidour: Voice AI's lower compute and data requirements enable small teams
“And in voice in particular, since the required compute is much lower and is that the same for data, really a few individuals can make stuff that is completely You know, just changing applications at very large scales.”