why aren't all 10 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Insight
Gu: High-quality speech synthesis requires multimodal foundation models
“And so actually to really get, like, perfect even just TTS or, like, speech-to-speech you actually really need to have, like, a model that has, More understanding, like at least of the language, but kind of like, it's not really an isolated component anymore. …”
Prediction Held up
Guo: Real-time generative audio and video apps will emerge within months
“I think we're going to see really, really cool experiences on the image video audio side, because as you say, if the models get smaller and they get better they also get, you know, and there's different architectures, like what Cartesia is working on, you're g…”
Assertion Supported
Gu: Optimal hybrid models use a 10:1 ratio of SSM to attention
“People have found that the optimal ratio tends to be mostly SSM layers with a little bit of attention. So maybe a ratio of like 10 to one, I know of at least Probably like five groups that have independently verified that this is kind of the optimal ratio of t…”
Disclosure
Goel: Cartesia intends to train an on-device multimodal model
“Yes. But, you know, we have our own sort of set of techniques that we're developing in order to be able to do that effectively. I think that I will maybe leave for another podcast. But I think, yeah, I think that is the intention at the end of the day is build…”
Insight
Gu: Aesthetic elegance was the primary driver behind inventing State Space Models
“People ask me, like, how do I treat my research problems, and my, I can't explain. My answer is just aesthetic. It's just like, there's something that I find elegant, and we're aesthetically pleasing about things, and to me, that's almost the most important th…”
Opinion
Goel: Current AI speech models fail to capture profession-specific vocal nuances
“And that's sort of the nuance that I don't think any models really capture well, which is like, you know, if you're a nurse, you need to talk in a different way than if you're a lawyer, or if you're a judge, or if you're a venture capitalist, you know, very di…”
Assertion Supported
Goel: Cartesia's Sonic TTS engine reduces typical voice latency by 150ms
“Even with what we've done with Sonic, we're already kind of shaving off, like, a 150 milliseconds off of you know, what they typically use.”
Disclosure
Goel: Cartesia aims to build edge-oriented infrastructure for SSMs
“I think in general for what we're trying to build is, is sort of the infrastructure to be able to train these models, make them run fast, and then bring them closer and closer to kind of be, you know, very edge oriented rather than cloud oriented.”
Assertion Supported
Goel: Cartesia runs Sonic TTS model locally on a standard Mac
“Yeah, I have a you know, our model running on our standard issue Mac here. Basically, this is you know, our text-to-speech model. Sonic on our playground is running in the cloud, and so you know, part of what I talked about earlier was how do you kind of bring…”
Disclosure
Goel: Cartesia aims to cut another 600ms of voice latency this year
“And so, you know, the roadmap is let's, let's get to the next 600 milliseconds and try to shape those off in the, over the course of the year.”