The Ledger, every show
Every statement that passed quotation and attribution checks, across all 44 shows. Pick shows below, then mix any filter with any other.
shows 




every show 44 of 44
Kumar: Near-term AI voice models will still lack dynamic conversational fluidity
“And the models that we have today, like CSM, for example, And probably some of the models that we'll have in the short term that will make the experience better will still not be modeling the conversational dynamics because they're making decisions that kind o…”
Kumar: Competitors will match Sesame's voice quality; there is no secret sauce
“The other companies, the other sort of chat products and so forth, they will get better voices. They're all, like, it's not gonna, we don't have some magical secret sauce on the technical side that is gonna be, like, impossible to replicate. They're gonna get …”
Kumar: Transcription-free conversational AI models are coming soon
“A pretty clear path that a lot of, I think, labs are taking, and we're taking as well, and will be in kind of future versions, is just kind of transcription free. Just go straight into the text component which will kind of obviate transcription entirely. That …”
Kumar: Open source won't solve AI voice and personality features
“But other parts in particular, kind of some of the personality aspects, some of the voice aspects, the speech generation, we didn't think and we still don't think will just be kind of done by the community. We think we will need to do it because that's kind of…”
Kumar: Sesame is actively developing smart glasses for its AI companions
“We mentioned on the website and we mentioned some of our launch content that we are working towards glasses as a form factor for Companions, or kind of this companion interface”
Kumar: Smartphones and laptops will not be replaced anytime soon
“No one's going to replace phones anytime soon, or laptops for that matter.”
Kumar: Sesame will build a unified audio-text transformer within months
“The path that we're going to take, I think, over the next few months is making a single transformer that does both audio understanding, content, text content generation, and speech generation.”
Kumar: Next Sesame AI models will feed audio natively into LLMs
“And so the kind of next versions of our models that will take audio and natively into the kind of LLM component will hopefully more and more pick up on those things.”
Kumar: Storytelling and AI will merge into new AI-native media categories
“And I think that we will see a lot of not just Sesame, but other kinds of media, let's say, that are sort of AI native in a way that Bring some creativity, bring some like storytelling into AI, or maybe bring AI into those categories. And I think they'll make …”
Kumar: Sesame achieves voice cloning via in-context learning prompt strings
“The model is this kind of, you know, it has kind of this in context learning style voice cloning. I mean, typically with some other kind of text-to-speech models, the voice cloning is kind of like an explicit feature. So it's sort of the model has dedicated ki…”
Kumar: No other open-source model generates multi-participant contextual audio
“At least to our knowledge, there's not another model out there that, that is open source that kind of is a sort of contextual thing where you kind of can put two participants in a conversation, even more, three, and generate kind of a conversation between them…”
Kumar: Speech research community will shift toward contextual AI architectures
“So, so the speech generation research community is very likely, I think, to move to more and more contextual architectures basically.”
Kumar: Sesame will not build a one-size-fits-all AI companion
“So we're certainly not going to, we don't see our product as like one companion that's the same for everyone. People have different preferences and that has to be a part of this kind of product category for sure.”
Kumar: Word error rate metrics for AI speech generation are now saturated
“Earlier on in the speech generation world in the community, very often you'd look at like word error rate where you look at transcription, like you kind of have a sentence and you generate and you transcribe it and you see if it's the same. And those metrics a…”
Kumar: Conversational AI will eventually rely on single models over heuristic pipelines
“I don't think you want to, in the long term, have those dynamics be like heuristics and so on, which they kind of are now. There are models involved in some heuristics and so forth. I think in the long term, it's just one model that is kind of naturally employ…”
Kumar: Transformers will remain the dominant AI sequence architecture short-term
“And I wouldn't bet against transformers, you know, not in the short term anyways.”
Kumar: Sesame will preserve AI companion personality as models improve
“They're making assistance. They're making utilities. I love those products. I use them all the time. They're great products. We want to make a companion. And so our prioritization of features and of, let's say, post training kind of personality, et cetera, wil…”
Kumar: Big tech companies will attempt to own conversational AI interface layer
“I think that over time, I think we will see more of these, you know, bigger companies trying to operate this layer. Like I said, I think that there is not enough effort on that right now, making these systems delightful to interact with, you know, and I think …”
Kumar: Developer plugins will be essential to future AI interfaces
“I think that the models still need to get better basically to utilize plugins essentially in a way that's kind of reliable enough that someone will go out and look for a plugin for, you know, their kind of downstream service of choice because they just want, y…”
Kumar: Sesame base model generates any voice with fine-tuning
“We are open sourcing the speech generation base model basically. And so the base model can generate any voice. It's quite conversational, but you do need to fine tune it probably if you want to get a particular personality or a particular kind of voice out of …”
Kumar: Sesame's speech generation is conditioned on full conversation audio
“The speech generation part is conditioned on the, on all the audio of the conversation.”
Kumar: Sesame targets sub-500 millisecond response times for voice AI
“We want You know, sub-five hundred millisecond response times, and a lot of things that feel like not a big deal, 50 milliseconds here, 50 milliseconds there, can really add up.”
Kumar: Larger speech models handle homographs and context-dependent pronunciation better
“And we see that as the models get bigger, they're much better at picking the right pronunciation in examples like this.”
Kumar: Sesame is developing a companion AI app with persistent memory
“We are making an app. We will make an app. I think for a little bit of time, it's going to still be kind of the demo experience. We want to support people using that for a long time, or, you know, we don't want to, we're not taking it away anytime soon from wh…”