Daily CEO Kwindla Hultman Kramer discusses the current landscape of voice models with Shawn Wang (swyx), contrasting experimental community models with enterprise speech transcription tools.
Prediction Not checkable as stated
Real-Time Video AI Will Inflect by Early 2026
“Like I really think real-time video is going to hit the same inflection point that voice did by the end of the year or early next year.”
Prediction Not checkable as stated
The Next TikTok Will Focus on Hyper-Personalized Interactive AI Video
“Well, I think we're going to have friends that are video in all our group chats, and the next TikTok is going to be not just hyper-personalized recorded content, but hyper-personalized interactive content.”
Assertion Not checkable as stated
99% of Current Monetizable Voice AI Use Cases Are Telephony
“99% of the monetizable voice AI use cases today are telephony.”
Prediction Not checkable as stated
Up to 75% of Future UX Interfaces Will Be Voice-Driven
“UX is going to be, you know, 50%, 60%, 75% voice in the future. I a hundred percent believe that, and I did not believe that, you know, two years ago, but the trend line is just really, I think, clear.”
Prediction Not checkable as stated
Real-Time Interactive AI Video Will Hit Consumers Before Enterprise
“I think real-time video may actually hit on the consumer side first. When it, when it's right, when it's on the right side of the uncanny valley, it's really, really compelling. And I think we're just starting to see some of that.”
Insight
Multimodal Real-Time Agents Require an Entirely New Programming Paradigm
“And it turns out that if you're building like agents that are multi-modal, multi-turn, real-time, it's just a totally different shape of programming problems and best practices than most other kinds of AI development even.”