Karan Goel, co-founder of Cartesia, discusses the low-latency performance of Cartesia's Sonic voice synthesis model compared to standard industry tools.
Prediction Not checkable as stated
Goel: AI models generating real-time data will replace traditional graphics engines
“In the future, like you will be essentially replacing, you know, graphics and rendering with essentially models that are outputting streams of data. In real time.”
Opinion
Goel: 3B parameter models are still too small and lack strong capabilities
“So the three V models are interesting, I think, but they aren't like, they're still small and not very capable.”
Opinion
Goel: Current AI speech models fail to capture profession-specific vocal nuances
“And that's sort of the nuance that I don't think any models really capture well, which is like, you know, if you're a nurse, you need to talk in a different way than if you're a lawyer, or if you're a judge, or if you're a venture capitalist, you know, very di…”
Disclosure
Goel: Cartesia intends to train an on-device multimodal model
“Yes. But, you know, we have our own sort of set of techniques that we're developing in order to be able to do that effectively. I think that I will maybe leave for another podcast. But I think, yeah, I think that is the intention at the end of the day is build…”
Disclosure
Goel: Cartesia aims to cut another 600ms of voice latency this year
“And so, you know, the roadmap is let's, let's get to the next 600 milliseconds and try to shape those off in the, over the course of the year.”
Disclosure
Goel: Cartesia aims to build edge-oriented infrastructure for SSMs
“I think in general for what we're trying to build is, is sort of the infrastructure to be able to train these models, make them run fast, and then bring them closer and closer to kind of be, you know, very edge oriented rather than cloud oriented.”