Multimodal AI
topic on 9 shows · 17 statements across 15 episodes
BG2 Pod
Latent Space
Lenny's Podcast
No Priors
Sourcery
the MAD Podcast
How I Built This
the a16z Podcast
Big Technology
17 statements about Multimodal AI, every show
Unified Transformer architectures simplified the training of natively multimodal AI models
“Even if this is not, like, the only architecture that would be, like, in a multi-model, but it made it really simple to train these models, like, natively, because you have, like, a single architecture and you can have all the modalities in, during training.”
Zhang: AI models will handle simple vision natively, using tools for complexity
“I think at least I want to bet on, you know, running their work natively together, the future for simple, I would say for simple or even intermediate difficult vision tasks. For example, kind of counting with less than 20 objects. I think for this kind of simp…”
Kaiser: Multimodal AI capabilities lag behind text performance
“The multimodal part still lags behind the text part to a large extent.”
Łukasz Kaiser: Robotics limitations reveal gaps in multimodal AI reasoning
“Robotics is probably just an illustration that we are not doing that well in multimodal and that we're not doing that well in general reasoning yet.”
Labenz: Non-text AI modalities will unify with language models over time
“We have seen this play out with text and image where you had your text only models and you had your image only models, and then they started to come together and now they've come really deeply together. And so I think you're going to see that across a lot of o…”
AI will reason over three-hour screen recordings in two to three years
“I think it's inevitable that in about two or three years, we will have models that are multimodal, large context windows, are capable of reasoning over, say, three hours of a screen recording.”
Gomez: Multimodal AI capabilities are table stakes for enterprise data
“But there's also like multimodal is Essential for understanding enterprise data, like PDF documents, where there's graphs and this type of thing, or understanding slide decks. A lot of the modalities that enterprises work in are visual. So it's sort of table s…”
Ulrich: Multimodal AI is uniquely valuable for financial invoice reconciliation
“When we think about things in financial services, big invoice reconciliation, bill pay, you'll have these things come in as PDFs, you'll have them come in as images, you'll have them come in as text in a variety of ways. As we start looking at things holistica…”
Winarski: Real-time visual AI assistants are likely, but privacy risks loom
“So more and more this, your vision, what you just said is not only possible, but likely unless people have access to your information and start sending you ads every which way.”
Winarski: Next AI Wave Will Move Beyond Language to Brainwaves and Robotics
“So you, we're going to start seeing generations beyond text and language going into other new breakthroughs, and we're going to see the same impact in those ways.”
Multimodal AI models will converge into unified single architectures
“So I think over time, I think the whole technology trend has been moving towards to a direction. A lot of all these things will be trained together. The multi-model model, multimedia, all get into one single model.”
Gerstner: Multimodal AI voice and video will disrupt BPO call centers
“I'll tell you a whole nother category of companies, Bill that have moved into that bucket and it's all these BPO companies, business process, outsourcing businesses, right? So, you know, because remember, it's not just about language now, right? BPO companies …”
Chaudhri: Multimodal AI eliminates the need for full computing screens
“Smartphones are great. They're really, really great at what they can do now, but they're very limited in terms of how much more they can do. And if you think about if they're able to leverage some of these multimodal inputs that use text and voice and sound an…”
Alex Immerman: Shared AI experiences will evolve from text to multimodal
“Looking ahead, these shared experiences are going to go from just being text to being multimodal.”
GitHub CPO: Sketch-to-code AI is for collaboration, not building production software
“I'm thinking about that as a better collaboration tool versus a production tool. Because right now, a lot of the time where we see challenges in articulating ideas is the communication. It's the clarity of thoughts. So if we can leverage AI to improve collabor…”
Gandhi: Multimodal AI will progress sequentially from text to audio-visual integration
“Multimodal, I think we're early stages now. So you're gonna see the early models, which are audio based. So I think about it like text, then images and text with Audio. And then you're kind of combining these medias together.”