Multimodal Models
topic on 7 shows · 13 statements across 11 episodes
Latent Space
Lenny's Podcast
the Neon Show
No Priors
the MAD Podcast
Big Technology
TBPN
13 statements about Multimodal Models, every show
Goel: Building Multimodal Models Lacks Any Established Recipe or Published Papers
“In a lot of these new areas, like multimodal models, like, there's no, Known recipe, right? Like you can't just go, go to the internet and say like, hey, this is how we're going to build the model. Here's, you know, here's the recipe we can follow. Here's a pa…”
Zeghidour: Large multimodal models are too massive to run voice profitably
“And at the same time, these models are so large, they cannot run at scale because they will just make everyone lose money in the process.”
LeCroix: Small OCR models are often cheaper than large multimodal models
“Sometimes it's a lot cheaper to use a small OCR model to just get the text that you care about and then potentially post-process it or deal with it with another system than to run it through a large multimodal model that will basically do the same thing but at…”
Kilpatrick: Multimodal Foundation Models Match or Beat Domain-Specific Vision Models
“Relative to today where you can literally just write a prompt and send images or videos to the model and have it do those tasks like with basically, you know, near or better accuracy than you would get from domain specific models is absolutely fascinating.”
Nguyen: Pixel-based perception is much harder to scale than language in AI
“Much of it is, like because right now the models operating on, like, pixels instead of, like, language or whatnot, like, pixels is actually really, really hard for the models because, like, perception or visual perception. I think there's still, like, a lot of…”
Tay: Multimodal AI architectures will eventually move completely to early fusion
“As early fusion models get more traction, I think the themes will start to get more and more, like, it's a bit like how all the tasks like unify, like from Like, two zero one nine to, like, now it's like all the tasks are unifying, now it's like all the modali…”
Gu: High-quality speech synthesis requires multimodal foundation models
“And so actually to really get, like, perfect even just TTS or, like, speech-to-speech you actually really need to have, like, a model that has, More understanding, like at least of the language, but kind of like, it's not really an isolated component anymore. …”
Goel: Cartesia intends to train an on-device multimodal model
“Yes. But, you know, we have our own sort of set of techniques that we're developing in order to be able to do that effectively. I think that I will maybe leave for another podcast. But I think, yeah, I think that is the intention at the end of the day is build…”
Albrecht: Vision is not essential for most coding and reasoning agent tasks
“And actually we found that for most of the kind of like code writing and reasoning problems that we care about, the visual part isn't really a huge important part of it.”
Luan: AI will converge into a universal byte model across all modalities
“Multimodal models are becoming more of a thing, we're behavioral cloning the visual world, but really what we're just going to have is this like universal byte model, right? Where like tokens of data that have high signal come in, and then all of those pattern…”
Multimodal models will completely supplant text-only large language models
“I actually think like it's really clear today. Multimodal models are the default foundation model, right? It's just going to supplant LLMs. Like why did you just train a giant multimodal model?”
Ruiz: Vision models effectively interpret mixed visual assets on infinite canvases
“The fun of the Infinite Canvas and Teal Draw in particular is that you could just dump like whatever you want onto the canvas. Screenshots, text, images, other websites sticky notes, all that stuff. And the model, even as something that was in preview, like th…”