why aren't all 17 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Insight
Kumar: Voice AI models must decide every 100 milliseconds for natural interaction
“You need to make decisions at the hundred millisecond, let's say, Time segment so that if you're talking and the other person, you know, starts sort of making some noises that make it seem like they're trying to interrupt you or they want to say something, you…”
Insight
Kumar: Theoretical gains won't unseat transformers without matching years of optimization
“But because of all the engineering work that the community has done around Transformers, it's like, you know, it's very good. And you're not going to just sort of unseat that, you know, just by an idea, right? There's a lot of work to be done.”
Insight
Kumar: AI voice interface success depends on product experience over model size
“We think that that interface layer, it's really not kind of a core bigger, bigger models, better, better reasoning question. It's really a product experience question. It's really a question of, can you make a system that people actually want to interact with,…”
Insight
Kumar: Voice cloning and prompting alone cannot create great AI personalities
“It takes more than just sort of, you know, voice clone plus change the prompt. Now you have a new character that's just as good as it would be if you spent a lot of time on it. It takes, I think today making a great personality voice interface system, we can't…”
Insight
Kumar: Multi-step AI agents need 99% reliability for daily adoption
“Doing, especially challenging, kind of multi-step things, you know, agents, as people say, I think to make that part of your everyday habits, it has to be like, 99%, you know, and right now, you know, every extra step the thing needs to take, There's some perc…”
Insight
Kumar: Builders underestimate current releases due to internal roadmap gaps
“When you build the thing, right, when you're building the product and using it every day, you know, there are some things that you work on that don't get into the demo because they're going to take longer and you want to ship the demo. You kind of know how big…”
Insight
Kumar: AI product optimization depends on hard-to-quantify qualitative user reactions
“But really, I think with some of these more product experience questions, there's something qualitative about it that is very hard to quantify. That is one of the big challenges internally, actually, is how do you hill climb effectively on what is really an ML…”
Insight
Kumar: Internal ML testing fails when teams exhaust fresh user reactions
“Misleading at times because you tried so much and you don't have, at least when we're trying it internally, you don't have such a diversity of users that you get kind of the first reaction over and over, right? You only get so many first reactions. And then wh…”
Insight
Kumar: Speech-to-text transcription misses essential non-verbal audio cues
“Humans, of course, convey a lot of information through Their speech that is not the words, the content of the speech and transcription misses that entirely.”
Insight
Kumar: Good ML taste means avoiding what APIs will soon commoditize
“I think from my perspective, good taste in ML today, because it's such a fast moving field with so many people working across, you know, open source and APIs and big labs and so forth. Really, you're trying to identify What part of the ecosystem or what part o…”
Insight
Kumar: Traditional text-to-speech sounds flat because non-neutral tones risk sounding inappropriate
“And that's probably why, or it's one of the reasons why historically voice assistants feel so flat is that traditional text of speech, it's kind of like it can only be flat. Or in other words, if it tries to not be flat, it's very likely wrong.”
Insight
Kumar: Seemingly easy side products like APIs create massive engineering drag
“Sometimes it feels like an API or something like that is like relatively easy to do. And, you know, it's not like, it's not maybe as hard as some of the other things that we're doing, but everything is a drag on engineering, right?”
Insight
Kumar: AI conversation is a distinct modality requiring core research
“I think that conversation, like human conversation, is kind of its own modality. And it is nowhere near done, right? There's so much more to do in the core research side to make it better.”
Insight
Kumar: True AI naturalness requires modeling turn-taking and backchannels
“I think to get these things to feel very, very natural and real, you do need to model the full conversation, the turn taking, the back channels, everything.”
Insight
Kumar: Early AI startups need flexible systems thinkers over niche specialists
“Especially when you're smaller, you know, you don't really want to harden, like, you know, you have this team that is super, super niche and doing only this thing, because you don't know exactly what the stack is going to look like tomorrow. Things change on t…”
Insight
Kumar: Achieving human realism in voice AI is harder than text
“I think you, I think it's much easier. It would be much easier to make a system that produces text chats with you that feels like you're texting a human because there's such a compression of like what the entity on the other side is into just like text. Wherea…”
Insight
Kumar: Adding generative modalities to AI models is harder than understanding
“It's much harder to add a modality to a pre-trained model than it is, add a generative modality, than it is to add an understanding modality.”