Everything Kwindla Hultman Kramer said on any show that made the record, most notable first. Each card names its show and opens the statement there.
Real-Time Video AI Will Inflect by Early 2026
“Like I really think real-time video is going to hit the same inflection point that voice did by the end of the year or early next year.”
The Next TikTok Will Focus on Hyper-Personalized Interactive AI Video
“Well, I think we're going to have friends that are video in all our group chats, and the next TikTok is going to be not just hyper-personalized recorded content, but hyper-personalized interactive content.”
99% of Current Monetizable Voice AI Use Cases Are Telephony
“99% of the monetizable voice AI use cases today are telephony.”
Up to 75% of Future UX Interfaces Will Be Voice-Driven
“UX is going to be, you know, 50%, 60%, 75% voice in the future. I a hundred percent believe that, and I did not believe that, you know, two years ago, but the trend line is just really, I think, clear.”
Real-Time Interactive AI Video Will Hit Consumers Before Enterprise
“I think real-time video may actually hit on the consumer side first. When it, when it's right, when it's on the right side of the uncanny valley, it's really, really compelling. And I think we're just starting to see some of that.”
Multimodal Real-Time Agents Require an Entirely New Programming Paradigm
“And it turns out that if you're building like agents that are multi-modal, multi-turn, real-time, it's just a totally different shape of programming problems and best practices than most other kinds of AI development even.”
Speech-to-Speech Models Are Not Yet Widely Used in Production
“What's happening with speech to speech models and APIs, super hot topic, not widely used yet in production.”
Enterprise B2B Unexpectedly Drove the Voice AI Monetization Inflection Point
“A little bit to everybody's surprise, the voice AI monetization pull actually came from enterprise and B to B use cases. So a little bit, everybody surprised the inflection point with voice AI in terms of monetizable use cases came on the business side. It's l…”
No State-of-the-Art Model Can Reason While Processing Continuous Input
“None of the SOTA models can kind of think while also taking input.”
NVIDIA's 600M Parameter Parakeet Model Tops Speech Transcription Leaderboards
“And then Parakeet is Nvidia's new speech model, speech transcription model. That's number one on the leaderboards. And it's just like very enterprise tuned, like really, really rock solid, reliable at fairly small number of weights, like six hundred million pa…”
Kyutai's Bidirectional Moshi Architecture Has Not Scaled to Large LLMs
“The, that Kyutai Moshi model you mentioned, which was my favorite academic paper last year, is a step towards a truly bi-directional streaming in both directions, thinking all the time, LLM. That work has not been kind of scaled up to, you know, large LLM size…”