Insight
Kumar: Voice AI models must decide every 100 milliseconds for natural interaction
“You need to make decisions at the hundred millisecond, let's say, Time segment so that if you're talking and the other person, you know, starts sort of making some noises that make it seem like they're trying to interrupt you or they want to say something, you…”
Prediction Not checkable as stated
Kumar: Near-term AI voice models will still lack dynamic conversational fluidity
“And the models that we have today, like CSM, for example, And probably some of the models that we'll have in the short term that will make the experience better will still not be modeling the conversational dynamics because they're making decisions that kind o…”
Insight
Kumar: Theoretical gains won't unseat transformers without matching years of optimization
“But because of all the engineering work that the community has done around Transformers, it's like, you know, it's very good. And you're not going to just sort of unseat that, you know, just by an idea, right? There's a lot of work to be done.”
Prediction Not checkable as stated
Kumar: Competitors will match Sesame's voice quality; there is no secret sauce
“The other companies, the other sort of chat products and so forth, they will get better voices. They're all, like, it's not gonna, we don't have some magical secret sauce on the technical side that is gonna be, like, impossible to replicate. They're gonna get …”
Insight
Kumar: AI voice interface success depends on product experience over model size
“We think that that interface layer, it's really not kind of a core bigger, bigger models, better, better reasoning question. It's really a product experience question. It's really a question of, can you make a system that people actually want to interact with,…”
Prediction Not checkable as stated
Kumar: Transcription-free conversational AI models are coming soon
“A pretty clear path that a lot of, I think, labs are taking, and we're taking as well, and will be in kind of future versions, is just kind of transcription free. Just go straight into the text component which will kind of obviate transcription entirely. That …”
Opinion
Kumar: Top AI labs under-invest in creative taste and humanities
“I do think that there is kind of an under-investment or an under-focus in the sort of Strong AI team world on product experience and sort of creative taste and kind of humanities maybe, in a sense, to kind of bring AI to experiences that kind of everyday peopl…”
Prediction Not checkable as stated
Kumar: Open source won't solve AI voice and personality features
“But other parts in particular, kind of some of the personality aspects, some of the voice aspects, the speech generation, we didn't think and we still don't think will just be kind of done by the community. We think we will need to do it because that's kind of…”
Prediction Open · timeframe Mar 2028
Kumar: Sesame is actively developing smart glasses for its AI companions
“We mentioned on the website and we mentioned some of our launch content that we are working towards glasses as a form factor for Companions, or kind of this companion interface”
Prediction Not checkable as stated
Kumar: Smartphones and laptops will not be replaced anytime soon
“No one's going to replace phones anytime soon, or laptops for that matter.”
Insight
Kumar: Voice cloning and prompting alone cannot create great AI personalities
“It takes more than just sort of, you know, voice clone plus change the prompt. Now you have a new character that's just as good as it would be if you spent a lot of time on it. It takes, I think today making a great personality voice interface system, we can't…”
Prediction Didn’t hold up
Kumar: Sesame will build a unified audio-text transformer within months
“The path that we're going to take, I think, over the next few months is making a single transformer that does both audio understanding, content, text content generation, and speech generation.”
Insight
Kumar: Multi-step AI agents need 99% reliability for daily adoption
“Doing, especially challenging, kind of multi-step things, you know, agents, as people say, I think to make that part of your everyday habits, it has to be like, 99%, you know, and right now, you know, every extra step the thing needs to take, There's some perc…”
Opinion
Kumar: Not enough AI teams are focused on the interface layer
“But I don't think there's enough companies and kind of teams working on the interface layer.”
Insight
Kumar: Builders underestimate current releases due to internal roadmap gaps
“When you build the thing, right, when you're building the product and using it every day, you know, there are some things that you work on that don't get into the demo because they're going to take longer and you want to ship the demo. You kind of know how big…”
Insight
Kumar: AI product optimization depends on hard-to-quantify qualitative user reactions
“But really, I think with some of these more product experience questions, there's something qualitative about it that is very hard to quantify. That is one of the big challenges internally, actually, is how do you hill climb effectively on what is really an ML…”
Insight
Kumar: Internal ML testing fails when teams exhaust fresh user reactions
“Misleading at times because you tried so much and you don't have, at least when we're trying it internally, you don't have such a diversity of users that you get kind of the first reaction over and over, right? You only get so many first reactions. And then wh…”
Insight
Kumar: Speech-to-text transcription misses essential non-verbal audio cues
“Humans, of course, convey a lot of information through Their speech that is not the words, the content of the speech and transcription misses that entirely.”
Prediction Open · timeframe Mar 2028
Kumar: Next Sesame AI models will feed audio natively into LLMs
“And so the kind of next versions of our models that will take audio and natively into the kind of LLM component will hopefully more and more pick up on those things.”
Disclosure
Kumar: Sesame trades complex AI reasoning for natural voice interaction
“So, you know, if you talk to Maya and Miles, you probably will not be able to get the same quality of like reasoning capabilities or intelligence as other As other systems, but in return, you're kind of getting this much more natural fluid interaction.”
Prediction Not checkable as stated
Kumar: Storytelling and AI will merge into new AI-native media categories
“And I think that we will see a lot of not just Sesame, but other kinds of media, let's say, that are sort of AI native in a way that Bring some creativity, bring some like storytelling into AI, or maybe bring AI into those categories. And I think they'll make …”
Disclosure
Kumar: Sesame built its open-source voice models from scratch
“Like we had to build the models that we're going to open source from scratch in order to get them to a point where they can achieve this experience.”
Insight
Kumar: Good ML taste means avoiding what APIs will soon commoditize
“I think from my perspective, good taste in ML today, because it's such a fast moving field with so many people working across, you know, open source and APIs and big labs and so forth. Really, you're trying to identify What part of the ecosystem or what part o…”
Disclosure
Kumar: Sesame is not building an API or developer-facing product
“We are not a developer facing business. We're not making an API.”