Multimodal AI Models
topic on 5 shows · 5 statements across 5 episodes
Latent Space
Lenny's Podcast
No Priors
the a16z Podcast
TBPN
5 statements about Multimodal AI Models, every show
Vahdat: Multimodal AI tools will become transformative productivity drivers within 12 months
“I think that what's going to happen in the next 12 months is the same thing is going to be happening with input and output of images and video to these models. And to the extent that even for images, Imagine them as productivity and educational tools, not just…”
Lord: Handshake engages hundreds of elite music students for AI multimodal data
“And we're engaging like thousands or not thousands, like probably hundreds of top music students at, you know, the weed music schools in the country who are improving models, understanding of music.”
Kilpatrick: AI models now outperform humans on most vision tasks
“If you look at like multimodal, like the fact that the models can like with better, better than just from a multimodal input perspective, better than humans are at like most vision tasks, like The number of products and like things that that unlocks is like tr…”
Wang: AI gets no positive transfer across modalities like video to text
“My understanding there's no positive transfer from learning in one modality to other modalities. So like training off of a bunch of video doesn't really help you that much with your text problems and vice versa.”
Tay: Vision models will unify screen intelligence and natural imagery without bifurcating
“I think at the end of the day, like, the models would become, like, I don't see that there will be, like, screen agents and, like, natural images. Humans, like, you can read what's on a screen, you can go out and appreciate the scenery, right? You're not, like…”