Ankit Kumar

Co-founder & CTO, Sesame · 1 appearance on the record.

computed by AI from the episodes · how this works → · full disclaimer →

founderexecutiveengineerscientist@ankitkumarsf ↗sesame.com ↗

Ankit Kumar is the co-founder and CTO of Sesame, where he leads engineering and research on conversational speech models. He previously co-founded Ubiquity6 and Pilot AI, and led the Clyde AI engineering team at Discord.

56statements → 25claims → 9claims resolved → 78%fully supported → 3.91/5average certainty → 1.95/5average debate potential →

7 supported 0 partly supported 2 contradicted 2 not yet assessed 14 not checkable as stated how the 25 claims stand · each chip opens the sources

17 predictions · 8 assertions · 2 opinions · 17 insights · 12 disclosures · every statement was checked. The predictions and assertions are the 25 claims: statements the public record can support or contradict. 9 are resolved, 2 are not yet assessed, and 14 name no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Ankit argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
Kumar: Sesame achieves voice cloning via in-context learning prompt strings
“The model is this kind of, you know, it has kind of this in context learning style voice cloning. I mean, typically with some other kind of text-to-speech models, the voice cloning is kind of like an explicit feature. So it's sort of the model has dedicated ki…”
Ankit Kumar Mar 15, 2025 ▶ 23:55 Building the Next Generation of Conversational AI

Their most notable contradicted claim

Prediction Didn’t hold up
Kumar: Sesame will build a unified audio-text transformer within months
“The path that we're going to take, I think, over the next few months is making a single transformer that does both audio understanding, content, text content generation, and speech generation.”
Ankit Kumar Mar 15, 2025 ▶ 1:00:20 Building the Next Generation of Conversational AI

Expressed certainty vs assessment result

none yet certainty 1
none yet certainty 2
0% certainty 3
100% certainty 4
100% certainty 5

weighted support: a fully supported claim counts one, a partly supported claim counts half. Each filled bar is clickable and opens exactly those claims; "none yet" means nothing said at that certainty level has resolved yet

How they sound: speaking style how? →

288 words/min while actually speaking · 4.6 um and uh per 1k words

No argument clarity score for Ankit Kumar: no usable question→answer exchanges on raw tape (a fair score needs 8+). We do not score a sample that small. Roundtable and news formats yield far fewer direct exchanges than interviews.

Measured by listening to 14,414 words across 1 episode, but every recording we have of Ankit Kumar is the aired feed, and an editor cleaned that audio before release. Some of the hesitation was cut before we ever heard it, so read these as floors: the true rates are at least this high. These are measurements of speaking style, not scores. How it's measured →

Everything Ankit Kumar said on the a16z Podcast that made the record, most notable first. Filter by type, assessment or year in the ledger →

Assertion Supported
Kumar: Sesame base model generates any voice with fine-tuning
“We are open sourcing the speech generation base model basically. And so the base model can generate any voice. It's quite conversational, but you do need to fine tune it probably if you want to get a particular personality or a particular kind of voice out of …”
Ankit Kumar Mar 15, 2025 ▶ 22:55 Building the Next Generation of Conversational AI
Assertion Not checkable as stated
Kumar: Sesame's speech generation is conditioned on full conversation audio
“The speech generation part is conditioned on the, on all the audio of the conversation.”
Ankit Kumar Mar 15, 2025 ▶ 29:18 Building the Next Generation of Conversational AI
Assertion Supported
Kumar: Sesame targets sub-500 millisecond response times for voice AI
“We want You know, sub-five hundred millisecond response times, and a lot of things that feel like not a big deal, 50 milliseconds here, 50 milliseconds there, can really add up.”
Ankit Kumar Mar 15, 2025 ▶ 45:15 Building the Next Generation of Conversational AI
Disclosure
Kumar: Sesame trained 1B, 3B, and 8B parameter models for speech generation
“So we published in the, in our blog post, we trained three variants. We trained 1,000,000,003 1,000,000,008 billion of just speech generation.”
Ankit Kumar Mar 15, 2025 ▶ 47:07 Building the Next Generation of Conversational AI
Assertion Supported
Kumar: Larger speech models handle homographs and context-dependent pronunciation better
“And we see that as the models get bigger, they're much better at picking the right pronunciation in examples like this.”
Ankit Kumar Mar 15, 2025 ▶ 48:31 Building the Next Generation of Conversational AI
Disclosure
Kumar: Sesame evaluates speech models against real human conversation continuations
“We also have some data sets that are kind of like Just two people in a conversation or sometimes they're actors, but it's trying to be a real conversation. And so we'll kind of take the conditioning of some snippet of the conversation and then show a human rat…”
Ankit Kumar Mar 15, 2025 ▶ 54:10 Building the Next Generation of Conversational AI
Prediction Held up
Kumar: Sesame is developing a companion AI app with persistent memory
“We are making an app. We will make an app. I think for a little bit of time, it's going to still be kind of the demo experience. We want to support people using that for a long time, or, you know, we don't want to, we're not taking it away anytime soon from wh…”
Ankit Kumar Mar 15, 2025 ▶ 59:04 Building the Next Generation of Conversational AI
Assertion Supported
Kumar: Sesame's AI voice companions currently cannot execute tasks
“Maya and Miles today, they can't do anything for you”
Ankit Kumar Mar 15, 2025 ▶ 1:26:00 Building the Next Generation of Conversational AI

Appearances (1)

EpisodeDateSpeaking time
Building the Next Generation of Conversational AI Mar 15, 2025 1h 9m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.