why aren't all 8 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Prediction Didn’t hold up
Kumar: Sesame will build a unified audio-text transformer within months
“The path that we're going to take, I think, over the next few months is making a single transformer that does both audio understanding, content, text content generation, and speech generation.”
Assertion Contradicted
Kumar: No other open-source model generates multi-participant contextual audio
“At least to our knowledge, there's not another model out there that, that is open source that kind of is a sort of contextual thing where you kind of can put two participants in a conversation, even more, three, and generate kind of a conversation between them…”
Assertion Supported
Kumar: Sesame achieves voice cloning via in-context learning prompt strings
“The model is this kind of, you know, it has kind of this in context learning style voice cloning. I mean, typically with some other kind of text-to-speech models, the voice cloning is kind of like an explicit feature. So it's sort of the model has dedicated ki…”
Prediction Held up
Kumar: Sesame will not build a one-size-fits-all AI companion
“So we're certainly not going to, we don't see our product as like one companion that's the same for everyone. People have different preferences and that has to be a part of this kind of product category for sure.”
Assertion Supported
Kumar: Sesame's AI voice companions currently cannot execute tasks
“Maya and Miles today, they can't do anything for you”
Assertion Supported
Kumar: Sesame base model generates any voice with fine-tuning
“We are open sourcing the speech generation base model basically. And so the base model can generate any voice. It's quite conversational, but you do need to fine tune it probably if you want to get a particular personality or a particular kind of voice out of …”
Prediction Held up
Kumar: Sesame is developing a companion AI app with persistent memory
“We are making an app. We will make an app. I think for a little bit of time, it's going to still be kind of the demo experience. We want to support people using that for a long time, or, you know, we don't want to, we're not taking it away anytime soon from wh…”
Assertion Supported
Kumar: Sesame targets sub-500 millisecond response times for voice AI
“We want You know, sub-five hundred millisecond response times, and a lot of things that feel like not a big deal, 50 milliseconds here, 50 milliseconds there, can really add up.”