Ankit Kumar, cofounder and CTO of Sesame, outlines the company's near-term research roadmap for conversational AI during a podcast with a16z General Partner Anjney Midha.
Insight
Kumar: Voice AI models must decide every 100 milliseconds for natural interaction
“You need to make decisions at the hundred millisecond, let's say, Time segment so that if you're talking and the other person, you know, starts sort of making some noises that make it seem like they're trying to interrupt you or they want to say something, you…”
Prediction Not checkable as stated
Kumar: Near-term AI voice models will still lack dynamic conversational fluidity
“And the models that we have today, like CSM, for example, And probably some of the models that we'll have in the short term that will make the experience better will still not be modeling the conversational dynamics because they're making decisions that kind o…”
Insight
Kumar: Theoretical gains won't unseat transformers without matching years of optimization
“But because of all the engineering work that the community has done around Transformers, it's like, you know, it's very good. And you're not going to just sort of unseat that, you know, just by an idea, right? There's a lot of work to be done.”
Prediction Not checkable as stated
Kumar: Competitors will match Sesame's voice quality; there is no secret sauce
“The other companies, the other sort of chat products and so forth, they will get better voices. They're all, like, it's not gonna, we don't have some magical secret sauce on the technical side that is gonna be, like, impossible to replicate. They're gonna get …”
Insight
Kumar: AI voice interface success depends on product experience over model size
“We think that that interface layer, it's really not kind of a core bigger, bigger models, better, better reasoning question. It's really a product experience question. It's really a question of, can you make a system that people actually want to interact with,…”
Prediction Not checkable as stated
Kumar: Transcription-free conversational AI models are coming soon
“A pretty clear path that a lot of, I think, labs are taking, and we're taking as well, and will be in kind of future versions, is just kind of transcription free. Just go straight into the text component which will kind of obviate transcription entirely. That …”