Ankit Kumar

Co-founder & CTO, Sesame · 1 appearance on the record.

computed by AI from the episodes · how this works → · full disclaimer →

founderexecutiveengineerscientist@ankitkumarsf ↗sesame.com ↗

Ankit Kumar is the co-founder and CTO of Sesame, where he leads engineering and research on conversational speech models. He previously co-founded Ubiquity6 and Pilot AI, and led the Clyde AI engineering team at Discord.

56statements → 25claims → 9claims resolved → 78%fully supported → 3.91/5average certainty → 1.95/5average debate potential →

7 supported 0 partly supported 2 contradicted 2 not yet assessed 14 not checkable as stated how the 25 claims stand · each chip opens the sources

17 predictions · 8 assertions · 2 opinions · 17 insights · 12 disclosures · every statement was checked. The predictions and assertions are the 25 claims: statements the public record can support or contradict. 9 are resolved, 2 are not yet assessed, and 14 name no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Ankit argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
Kumar: Sesame achieves voice cloning via in-context learning prompt strings
“The model is this kind of, you know, it has kind of this in context learning style voice cloning. I mean, typically with some other kind of text-to-speech models, the voice cloning is kind of like an explicit feature. So it's sort of the model has dedicated ki…”
Ankit Kumar Mar 15, 2025 ▶ 23:55 Building the Next Generation of Conversational AI

Their most notable contradicted claim

Prediction Didn’t hold up
Kumar: Sesame will build a unified audio-text transformer within months
“The path that we're going to take, I think, over the next few months is making a single transformer that does both audio understanding, content, text content generation, and speech generation.”
Ankit Kumar Mar 15, 2025 ▶ 1:00:20 Building the Next Generation of Conversational AI

Expressed certainty vs assessment result

none yet certainty 1
none yet certainty 2
0% certainty 3
100% certainty 4
100% certainty 5

weighted support: a fully supported claim counts one, a partly supported claim counts half. Each filled bar is clickable and opens exactly those claims; "none yet" means nothing said at that certainty level has resolved yet

How they sound: speaking style how? →

288 words/min while actually speaking · 4.6 um and uh per 1k words

No argument clarity score for Ankit Kumar: no usable question→answer exchanges on raw tape (a fair score needs 8+). We do not score a sample that small. Roundtable and news formats yield far fewer direct exchanges than interviews.

Measured by listening to 14,414 words across 1 episode, but every recording we have of Ankit Kumar is the aired feed, and an editor cleaned that audio before release. Some of the hesitation was cut before we ever heard it, so read these as floors: the true rates are at least this high. These are measurements of speaking style, not scores. How it's measured →

Everything Ankit Kumar said on the a16z Podcast that made the record, most notable first. Filter by type, assessment or year in the ledger →

Disclosure
Kumar: Sesame is open sourcing its speech model, not the full demo
“We're not open sourcing the demo. We're open sourcing the speech generation model that is powering the voice of the demo.”
Ankit Kumar Mar 15, 2025 ▶ 20:58 Building the Next Generation of Conversational AI
Assertion Supported
Kumar: Sesame achieves voice cloning via in-context learning prompt strings
“The model is this kind of, you know, it has kind of this in context learning style voice cloning. I mean, typically with some other kind of text-to-speech models, the voice cloning is kind of like an explicit feature. So it's sort of the model has dedicated ki…”
Ankit Kumar Mar 15, 2025 ▶ 23:55 Building the Next Generation of Conversational AI
Assertion Contradicted
Kumar: No other open-source model generates multi-participant contextual audio
“At least to our knowledge, there's not another model out there that, that is open source that kind of is a sort of contextual thing where you kind of can put two participants in a conversation, even more, three, and generate kind of a conversation between them…”
Ankit Kumar Mar 15, 2025 ▶ 25:47 Building the Next Generation of Conversational AI
Insight
Kumar: Traditional text-to-speech sounds flat because non-neutral tones risk sounding inappropriate
“And that's probably why, or it's one of the reasons why historically voice assistants feel so flat is that traditional text of speech, it's kind of like it can only be flat. Or in other words, if it tries to not be flat, it's very likely wrong.”
Ankit Kumar Mar 15, 2025 ▶ 28:42 Building the Next Generation of Conversational AI
Prediction Not checkable as stated
Kumar: Speech research community will shift toward contextual AI architectures
“So, so the speech generation research community is very likely, I think, to move to more and more contextual architectures basically.”
Ankit Kumar Mar 15, 2025 ▶ 28:56 Building the Next Generation of Conversational AI
Insight
Kumar: Seemingly easy side products like APIs create massive engineering drag
“Sometimes it feels like an API or something like that is like relatively easy to do. And, you know, it's not like, it's not maybe as hard as some of the other things that we're doing, but everything is a drag on engineering, right?”
Ankit Kumar Mar 15, 2025 ▶ 37:43 Building the Next Generation of Conversational AI
Prediction Held up
Kumar: Sesame will not build a one-size-fits-all AI companion
“So we're certainly not going to, we don't see our product as like one companion that's the same for everyone. People have different preferences and that has to be a part of this kind of product category for sure.”
Ankit Kumar Mar 15, 2025 ▶ 38:37 Building the Next Generation of Conversational AI
Insight
Kumar: AI conversation is a distinct modality requiring core research
“I think that conversation, like human conversation, is kind of its own modality. And it is nowhere near done, right? There's so much more to do in the core research side to make it better.”
Ankit Kumar Mar 15, 2025 ▶ 40:05 Building the Next Generation of Conversational AI
Insight
Kumar: True AI naturalness requires modeling turn-taking and backchannels
“I think to get these things to feel very, very natural and real, you do need to model the full conversation, the turn taking, the back channels, everything.”
Ankit Kumar Mar 15, 2025 ▶ 42:38 Building the Next Generation of Conversational AI
Insight
Kumar: Early AI startups need flexible systems thinkers over niche specialists
“Especially when you're smaller, you know, you don't really want to harden, like, you know, you have this team that is super, super niche and doing only this thing, because you don't know exactly what the stack is going to look like tomorrow. Things change on t…”
Ankit Kumar Mar 15, 2025 ▶ 46:08 Building the Next Generation of Conversational AI
Assertion Not checkable as stated
Kumar: Word error rate metrics for AI speech generation are now saturated
“Earlier on in the speech generation world in the community, very often you'd look at like word error rate where you look at transcription, like you kind of have a sentence and you generate and you transcribe it and you see if it's the same. And those metrics a…”
Ankit Kumar Mar 15, 2025 ▶ 52:51 Building the Next Generation of Conversational AI
Insight
Kumar: Achieving human realism in voice AI is harder than text
“I think you, I think it's much easier. It would be much easier to make a system that produces text chats with you that feels like you're texting a human because there's such a compression of like what the entity on the other side is into just like text. Wherea…”
Ankit Kumar Mar 15, 2025 ▶ 56:40 Building the Next Generation of Conversational AI
Insight
Kumar: Adding generative modalities to AI models is harder than understanding
“It's much harder to add a modality to a pre-trained model than it is, add a generative modality, than it is to add an understanding modality.”
Ankit Kumar Mar 15, 2025 ▶ 1:00:30 Building the Next Generation of Conversational AI
Prediction Not checkable as stated
Kumar: Conversational AI will eventually rely on single models over heuristic pipelines
“I don't think you want to, in the long term, have those dynamics be like heuristics and so on, which they kind of are now. There are models involved in some heuristics and so forth. I think in the long term, it's just one model that is kind of naturally employ…”
Ankit Kumar Mar 15, 2025 ▶ 1:03:36 Building the Next Generation of Conversational AI
Disclosure
Kumar: Sesame is developing diffusion-based audio generation models
“We are also working, by the way, on kind of ideas that make the audio generation part diffusion.”
Ankit Kumar Mar 15, 2025 ▶ 1:06:57 Building the Next Generation of Conversational AI
Prediction Not checkable as stated
Kumar: Transformers will remain the dominant AI sequence architecture short-term
“And I wouldn't bet against transformers, you know, not in the short term anyways.”
Ankit Kumar Mar 15, 2025 ▶ 1:08:15 Building the Next Generation of Conversational AI
Disclosure
Kumar: Sesame intentionally includes speech imperfections to make AI sound natural
“My and Miles, they might sort of say the wrong thing or kind of like back up a little bit and say something else or something. And that's on purpose, of course.”
Ankit Kumar Mar 15, 2025 ▶ 1:11:02 Building the Next Generation of Conversational AI
Prediction Not checkable as stated
Kumar: Sesame will preserve AI companion personality as models improve
“They're making assistance. They're making utilities. I love those products. I use them all the time. They're great products. We want to make a companion. And so our prioritization of features and of, let's say, post training kind of personality, et cetera, wil…”
Ankit Kumar Mar 15, 2025 ▶ 1:14:19 Building the Next Generation of Conversational AI
Prediction Not checkable as stated
Kumar: Big tech companies will attempt to own conversational AI interface layer
“I think that over time, I think we will see more of these, you know, bigger companies trying to operate this layer. Like I said, I think that there is not enough effort on that right now, making these systems delightful to interact with, you know, and I think …”
Ankit Kumar Mar 15, 2025 ▶ 1:30:09 Building the Next Generation of Conversational AI
Disclosure
Kumar: The ChatGPT plugin system I built failed to fully take off
“I did this chat to be plugin system before, and I think it's still there probably. And it kind of didn't fully take off really.”
Ankit Kumar Mar 15, 2025 ▶ 1:32:29 Building the Next Generation of Conversational AI
Prediction Not checkable as stated
Kumar: Developer plugins will be essential to future AI interfaces
“I think that the models still need to get better basically to utilize plugins essentially in a way that's kind of reliable enough that someone will go out and look for a plugin for, you know, their kind of downstream service of choice because they just want, y…”
Ankit Kumar Mar 15, 2025 ▶ 1:32:53 Building the Next Generation of Conversational AI
Disclosure
Kumar: Sesame's current demo cannot detect user emotional tone
“The current demo does not sort of hear the user from the perspective of their paralinguistic kind of emotional tone and so forth.”
Ankit Kumar Mar 15, 2025 ▶ 7:06 Building the Next Generation of Conversational AI
Disclosure
Kumar: Sesame's entire software and ML team is under 15 people
“The full software team today is still under 15 people, and so we just don't, that's including ML and infrastructure and everything.”
Ankit Kumar Mar 15, 2025 ▶ 8:28 Building the Next Generation of Conversational AI
Disclosure
Kumar: Sesame is not pre-training frontier LLMs at scale
“You know, we are not A frontier model company. We're not pre-training LLMs at insane scale and so forth.”
Ankit Kumar Mar 15, 2025 ▶ 9:40 Building the Next Generation of Conversational AI

Show 8statements(8 left)

Appearances (1)

EpisodeDateSpeaking time
Building the Next Generation of Conversational AI Mar 15, 2025 1h 9m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.