The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Neil Zeghidour argument clarity score 4.2/5 from 14 exchanges on raw tape · average scores: directness 4.6 · coherence 4.2 · precision 4 · compression 3.8 record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
14exchanges match
14on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q And, uh, voice cloning, is that a, is that a use case?

A Actually with Gradium, we have the best of, um, of the, of the industry. And the best means not only replicating like the specific characteristic of someone, but I mean, also the accent, some unusual recording condition. And so in a lot of contexts, for example, if you want to create a vintage sounding character with old radio effects, and it's something that we do pretty well. If you want to have A robot voice. We can do that pretty well as well. If you want to have any kind of accent or speaking style like posh or, uh, more, uh, you know, laid back or more urban, any kind of social aspects of, uh, and individual aspect of voice is something that we, uh, that we replicate pretty well. I think what is interesting is that, uh, voice learning in itself where I see the most potential is about Uh, creative, creating interactive experiences around, uh, licenses. I tried to pitch, for example, you know, who wants to be a millionaire? Uh, there is a video game, uh, and the questions can be generated on the fly. You would like the host to, the voice of the host to pull on them all the time. So this one makes a lot of sense with cloning a specific voice because you want to replicate the voice of a character, of a person, of an athlete, of a K-pop star, whatever, I don't know, you know, any kind of experience where people want to engage with a voice that they know. But they want to go th…

AI assessment note: “voice learning in itself where I see the most potential is about Uh, creative”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q obviously you are a new entrant in the field of AI, where there's tremendous amounts of competition. So I think for voice in particular, like the obvious question is why has OpenAI or Google or Meta not already won voice AI? And I think you probably alluded to Some of the reasons up front, but, uh, what, why is that? Why can a small company hope to become the leader?

A So one thing I mentioned was, uh, if you have the right team, it can be extremely small and still make a significant impact. Other, uh, arguments I think is, one is focus. So, for example, if you look at large multimodal models, right? Like these generic models that understand images and can generate text and can produce code and so on. You have, uh, like a limited budget, which is the number of parameters and data you're going to feed to your model. When you want to add speech to them, you're fighting with coding and, uh, image understanding and so on. So you are playing with a lot of trade-offs that are irrelevant to the task that you want to solve. And at the same time, these models are so large, they cannot run at scale because they will just make everyone lose money in the process. The only A format that makes sense for speech models to run at scale is to be extremely compact, which also means that the training resources you need to train them are much smaller than what you need in, uh, to train other kinds of models. So the resources are not as challenging as for text models. I think also another aspect is, uh, in a way, not trying to just make a conversational product. So really making building blocks so that people can build the product. So we could make The Gradium conversational assistant and think a lot about its capabilities and what it can do and what it cannot do …

AI assessment note: “if you have the right team, it can be extremely small and still make a significant impact.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q You mentioned data a second ago. How does that work for voice AI, and how does that compare to text AI? Obviously, text AI, the LLMs are training on the whole internet, but presumably there's a lot less Audio and speech data to train on? How does that work?

A So if you do the math, basically like training on a few trillions of tokens, which is what you will do for, you know, like a basic text model that will amount to hundreds of millions of hours of speech or something like that, which is kind of amounts that are very hard to get. I think this is a very interesting question that comes up in a lot of discussions and everybody has their theories. In particular, one impact, uh, one, let's say, one attribute of speech data Was that if you train a conversational model on speech data, it's going to be much less intelligent than a model trained on text. And I think it's because when you listen to speech data, it's, uh, the density of information is, uh, is much lower than you would have in text. So you don't have Wikipedia, like, uh, you know, in speech data, you don't have Stack Overflow, Reddit, and so on. Getting your model to learn about the world from speech, I think it's a terrible idea. I think you should start. I mean, you know, we have text on that. So what we did for mostly is we started from the text model and then we, we, we, we, We took this text model and, and, and trained it on speech while trying to prevent as much as possible a loss of intelligence. So all the time we will recompute the text metrics and they will degrade, but we are trying, you know, like to keep it a bit contained, but indeed the, the quality of, uh, of …

AI assessment note: “density of information is, uh, is much lower than you would have in text”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q And then there's, uh, an additional aspect to this, which is that voice can also, and should be pretty often on device versus, uh, an API call to the cloud. Is that, is that fair?

A I think what is very, very challenging right now is if you want to have the full intelligence on device, like the, like your full conversational AI on device. Honestly, I would say at this point, if you want such a model to be useful, we are not there yet, right? You can have a Something that can chit chat a bit and it will be decent. Or we also have shipped, uh, models on device, but they are much more constrained in terms of applications. So for example, we, we started a year ago with, uh, uh, on device speech to speech translation, which is something that makes a lot of sense because when you're traveling, maybe, you know, you don't have a data plan that is, uh, going in every country. So it makes sense to have something that works on your phone if you want to order at a restaurant, something like that. I think it's a particularly adapted use case, but now we're also, Uh, we released two weeks ago a model called Pocket ETS that not only is on device, but CPU only. So, uh, there are already mods for AAA video games, uh, where the NPCs can be powered through these voice models. And now you unlock a completely new kind of, uh, of applications because on device models allow to do very large scale personalized content, uh, that will be, uh, economically not realistic with an API. So again, these kind of things is, uh, You know, if you want to make meaningful progress in that dire…

AI assessment note: “on device models allow to do very large scale personalized content, uh, that will be, uh, economically not realistic with an API”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q Amazing. Amazing. You played an incredibly, uh, important pioneering, uh, role in, I guess, the current state of voice areas. And then what was the next stop after that?

A At the time where, where, uh, uh, Gemini started at, uh, at Google, that's where, that's when I, I left. So I wanted to, to, you know, to create a small research environment that reminded me of the early days of fair or Google brain. So very small team, elite, uh, no distraction. Sorry to say that no product manager or no, you know, like just, just research scientists, no emails that just, just, you know, locked in a room with the machines and, uh, focused on, on science. And in particular, what the goal for me was really to keep Working on fundamental research and keep pushing the field forward and training students and so on, because I felt very grateful to have been able to do research in such an open environment. It was also obvious for me that, and for this, I think I agree 100% with Yann Lequin, the fact that what made AI dynamic and get from ImageNet in 2012 to where we are right now today, Is open research. That's, uh, because it's kind of a worldwide collaboration and everybody benefits from the progress of everyone. So that's, for me, it was important for the field itself to, to keep this, uh, going. And so we decided to create a nonprofit with the help of, uh, of Eric Schmidt, uh, Xavier and Rodolf Sadeh. So for the anecdote, the code name was Sphere because that's the name of the restaurant where we discussed the project. And, uh, we then understood that we could ne…

AI assessment note: “and so that's how we created Qtai.”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q You mentioned benchmarks a minute ago, and that seems to be a really interesting question for voice and video as well, and images. How do you make a case that your technology is better than the next provider? Because some of it seems to be a little bit around vibes, right? Like how you feel When you're on the receiving end of a voice AI.

A One thing that is, uh, that is clear is that you can only trust human judgment. People have tried to make objective proxies of, uh, human judgment. Like that would be a neural network that listens to an audio and gives it a grade. Uh, it sucks. Like so many people try and it works on their constrained setting and on real audio it doesn't work at all and completely breaks. So we don't trust anything about, uh, No, but our ears. So we do a lot of blind tests internally. We do a lot of blind tests externally. So we are working with human judgment constantly. So every single decision we make is based on human listening. Uh, we don't trust metrics at all. So it's fundamentally subjective experience, the quality of audio, but there are some things that are going to be widely shared. Reserves like the, the, the prosody, so the, the tone and the rhythm are natural or not. A lot of people would agree on that. Is a voice nice or not? Nobody agrees on that. And so then the only way for me to claim to have, uh, the best solution is to have the largest, uh, catalog and most diverse set of voices that people can pick. Because then, you know, it's, uh, the kind of what voice are people going to like? This really depends on, uh, Uh, between people. I had faced that in the past when we did Music LM at Google. So for the first time we made, you know, text to music. So you could type like, uh, uh…

AI assessment note: “the only way for me to claim to have, uh, the best solution”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q use cases? And the, you know, there's certainly an argument, um, that you'll hear a lot of people saying voice is great, but like most of the time, like I'm, At the office, like, the last thing I want is, like, for people to hear my conversation, and therefore I don't want to talk to a machine. So where, where does voice fit in, um, that vision of the future?

A So for example, I used to think that one obvious application where voice was kind of irrelevant was coding, because it's fundamentally, you're not going to read code out loud, right? Yet now, since coding is going more and more towards vibe coding, which is natural language, it makes a lot of sense to do it by, by speaking. And, you know, now people are developing products that allow you to, you know, dispatch orders to coding agents, uh, in a way that is much more efficient than if you had to type in each different window to each of them. Even prompting LLMs now is, uh, doing by voice is much more convenient rather than typing. I still agree that there is one part which is more social about Uh, what the offices environment will look like? I don't know. Maybe we will just also rethink, uh, the way we just structure office environments. What is sure is, uh, now people have AI assistants that are almost colleagues, right? Uh, I mean, you talk to any software engineer, the anthropomorphization of cloud code is, I find it extremely funny. You know, even the, the, the verb coding, I mean, it's going to be coding pretty soon. And so these people, you know, they will, If it's more convenient to interact with their main, uh, tool through voice, uh, that will justify also rethinking office spaces, I guess. So, so, uh, yeah, I think there will be workarounds that, and we will naturally f…

AI assessment note: “now people are developing products that allow you to, you know, dispatch orders to coding agents”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q The obvious next question is, uh, if, if nobody can agree on whether this model is better than that model, then, uh, aren't all the models more or less the same and therefore the entire voice AI model industry is sort of like commoditized?

A All the models being, uh, the same. I think then people talk about TTS already, right? Which is much more constrained than voice AI. So In text to speech, I mean, there are factual metrics about accuracy and latency, and then there are more subjective things around expressivity and so on. But you can, again, you can really make a difference by making it more controllable, more customizable, and so on and so forth. So I don't think that it's clear that the best TTS, the most controllable, the most robust, the most smart in terms of expression is in front of us. Nothing is, is close to it yet. And then there is everything that is not TTS. Uh, you look at transcription. Now again, what if a lot of people are speaking? So I was talking about diarization as the least sexy problem ever. At the same time, it's, uh, it's a extremely useful problem. And you look at the error rates and they are, they are very bad. I mean, it's still just not working in difficult cases where you have a podcast with a lot of people talking at the same time, just completely breaks. Full duplex. We, you know, we did Moshi a year and a half ago. Still nobody has made it into a product. There are so many things that are just not existing today that I think the, the communitization, maybe it will happen someday, but we are very, very far from it, honestly. And the gap in, uh, you know, there is already bridging…

AI assessment note: “the communitization, maybe it will happen someday, but we are very, very far from it”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q And what about a hardware? So you, you mentioned earlier, this is not a big hardware. GPU play compared to the big LLMs, is that, is that correct?

A If you want to have voice models to run at scale, and it makes sense from a, you know, in terms of economics, you need to have this model compact anyway. If you think about having NPCs in the game, where people are going to pay, I don't know, 70 bucks or 90 bucks or whatever for a game, and then they want to talk to it three hours a day, because it's a very good game, so you spend a lot of time on it. You can, if you have a large model, it's just impossible. Not only it doesn't fit on the GPU, but even for APIs, it just wouldn't make sense economically. So fundamentally, I think these models need to be small, but sometimes they require access to large models because they are solving a complex task. I'm quite state, in a way stating the obvious, but selective and adaptive compute usage based on the context and the difficulty of the task at hand, the task being as precise as Like the next few words. Can I just answer like that? Because somebody say hey, and I just say hey, they ask for like the next flight to SF, and I have to look up on internet to find the time. So being very selective about when to compute, to use compute, I think that's, that's the only way to, for all of it to make sense economically.

AI assessment note: “fundamentally, I think these models need to be small”

Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q So just a few days ago, uh, there was this big announcement by Alibaba slash Quen that they were open sourcing the Quen III TTS family for voice design, clone, and generation. How do you think about, uh, open source, uh, in your world as Gradium? Is that a friend? Is that a foe?

A So if you look, if you read the paper, you will find our, uh, our names, uh, in several pages. It's mostly, uh, inspired from Uh, the Moshi architecture, like pretty much every, uh, uh, model right now, even the Voxstral model that was released by Mistral, uh, two weeks ago is also, uh, based on, uh, on our, uh, framework. I think that's really interesting because this proves that, uh, uh, there are things that we, you know, that we do right because everybody is building on them. At the same time, I would say it's quite an advantage because I would say not, there are two kinds of paper research papers. There are research papers that are meant to be as explicit and reproducible as possible, which is what we try to do when we do one. And there are some that are more about, I would say, uh, marketing in the sense that they are mostly focused on the results and the performance rather than explaining the underlying mechanism and the data and so on and so forth. So nice thing of build, people building around the frameworks we introduced is that even when they don't give details, we can infer the details. So in a way, in a competitive landscape, I think it's, uh, It's quite an advantage because, in a way, it would be more challenging if people were transitioning to something completely different from what we've been doing, because, you know, then we will could not infer anything from …

AI assessment note: “in a competitive landscape, I think it's, uh, It's quite an advantage”

Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q So we'll get into some of the technical details in a minute, but, um, at a high level, why, uh, has voice AI being, I guess, the most Underdeveloped modality. There's been, obviously, extraordinary progress on text AI, and then image AI, and then video AI, but it seems that voice has been a little bit the, the poor parent in terms of progress. Why is that?

A I don't want to be mean to my people, the speech scientists, but historically, for some reason, uh, voice did not attract, like, the visionaries in, in machine learning, right? So even if you looked at the dynamics in conferences, Uh, if you proposed a new method, uh, like fundamental algorithm, and you wanted it to be accepted in a prestigious venue, you had to have an application either in computer vision, like image classification or NLP. If you did it in speech, you would get rejected because it was like, uh, to speech. And at the same time, the, the prestige of speech conferences, uh, used to be much lower than that of, uh, computer vision or NLP. So honestly, I don't really know why, because when you look at the details, the first big success of deep learning, everybody knows the AlexNet model in, uh, uh, where for the first time, uh, you know, you had a deep learning model, uh, outperforming every single alternative on image classification. But actually the really first big success in deep learning was speech recognition. We've worked from Geoffrey Hinton himself.

AI assessment note: “voice did not attract, like, the visionaries in, in machine learning”

Answered raw tape D 5 · C 4 · P 3 · Cm 3 3.90

Q Fascinating. To put it in numbers, how many would you say people there are in the world with that expertise? Are we talking about? 100? 500? 10?

A Between 10 and 100? No, I would say. 50? I don't know. It's hard to say. But, uh, yeah, I think it's very few and, uh, really meaningful contributions that have pushed the field forward have been made by very small groups of people. And I think that's Also what's nice. So AI, I think in general is, is one field where individuals can have a disproportionate impact because, you know, the amount of things you can, you can do by yourself, you have access to, to compute and, uh, and data sets is huge. And in voice in particular, since the required compute is much lower and is that the same for data, really a few individuals can make, uh, stuff that is completely, uh, You know, just changing, uh, applications at very large scales.

AI assessment note: “Between 10 and 100? No, I would say. 50?”

Answered raw tape D 4 · C 4 · P 4 · Cm 3 3.85

Q What happens behind the scenes? When you have that noisy environment, like what, what does the model actually do? And how do you solve that problem? Like if you're calling a receptionist in a restaurant and like he or she picks up and there's super light in the back, how do you solve that problem?

A One thing that is, uh, really important is to have, uh, uh, several, uh, microphones. So we do, we have two ears and it's extremely useful because that's the lowest, for example, to localize An audio source. The reason, the why, the reason why we can localize where our sounds come from, because we have a small time difference between the time when it arrives in this year and this year. So if it arrives here before, my brain will understand it will come this and you get like a longer and then the acceleration of the phase gives you the other dimension. And so that's how you can locate in three. So having the ability of, uh, of doing this specialization, understanding, you know, where the sources are coming from and which one is saying what and so on. Again, it's, it's both a hardware and software issue. Creating training data for that is extremely challenging. Honestly, the level of robustness that we have as humans on that is, uh, is extremely high. Also, you know, we are all lip reading. We don't know that we are doing it, but everybody is lip reading all the time. That helps understanding a lot in noisy environments in particular. We do it unconsciously, but we all do it. I think interactions on the phone are Quite okay because, you know, the mouth is close to the phone and the phone can do a lot of work to enhance the quality. But now, you know, people want to have a robot t…

AI assessment note: “One thing that is, uh, really important is to have, uh, uh, several, uh, microphones.”

Answered raw tape D 3 · C 3 · P 4 · Cm 3 3.25

Q uh, uh, accounts, um, sort of misfired the other day, uh, by, uh, saying that he was going to allocate, uh, thirty million euros to AI when In reality, it meant a specific program to attract like a few academics to, to, to, to France. But, um, so what, on the ground there, what, what is your sense of the current state and the strengths and weaknesses of the friends?

A Yeah, a lot of things to say about that. I was born and raised in Paris. I did all my career there. Uh, Facebook arrives when I started my PhD and then Google Brain moved to Paris and Google DeepMind afterwards. I'm what we can call terminally online in the sense that I love The very mean memes against Europe. So there's this guy, I don't remember his name, it's like a fake Swedish name, and he keeps posting about how, you know, he has, after only, uh, 20 meetings, he has contributed a 10,000 euros check. And it's a compliance first company and everything, you know, but I love it. It's very mean. Honestly, I love it because it's very mean. I'm like, I love when people have so much time to spend just to be mean. I mean, I think it's, it's a quite, uh, you know, I respect that, but, uh, At the same time, it's so far from, uh, from reality. So, you know, uh, French AI and European AI. So European AI is mostly French AI, to be fair. Uh, there is also Germany, but a lot of it is in France. And before French companies, as I was French talent in American companies. So again, a lot of the current audio generative models of Google, a global company, obviously, uh, were developed between Paris and Zurich. Uh, actually most of it. Uh, Lama was started in Paris. Dino, which is the, uh, most groundbreaking vision work from Facebook was developed in Paris. A lot of things have been developed…

AI assessment note: “Lama was started in Paris. Dino, which is the, uh, most groundbreaking vision work”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.