Multimodal Models

topic on 7 shows · 13 statements across 11 episodes

Latent Space Lenny's Podcast the Neon Show No Priors the MAD Podcast Big Technology TBPN

13 statements about Multimodal Models, every show

NEON SHOW Insight
Goel: Building Multimodal Models Lacks Any Established Recipe or Published Papers
“In a lot of these new areas, like multimodal models, like, there's no, Known recipe, right? Like you can't just go, go to the internet and say like, hey, this is how we're going to build the model. Here's, you know, here's the recipe we can follow. Here's a pa…”
Karan Goel Aug 7, 2026 ▶ 15:56 The Billion Dollar AI Lab Founder Who Sees The Future First | Karan Goel, Founder & CEO of Cartesia
MAD Assertion Not checkable as stated
Zeghidour: Large multimodal models are too massive to run voice profitably
“And at the same time, these models are so large, they cannot run at scale because they will just make everyone lose money in the process.”
Neil Zeghidour Feb 19, 2026 ▶ 32:37 Voice AI’s Big Moment: Top Researcher on Why Everything Is Changing (Neil Zeghidour, Gradium AI)
MAD Insight
LeCroix: Small OCR models are often cheaper than large multimodal models
“Sometimes it's a lot cheaper to use a small OCR model to just get the text that you care about and then potentially post-process it or deal with it with another system than to run it through a large multimodal model that will basically do the same thing but at…”
Timothée LeCroix Feb 12, 2026 ▶ 49:25 Mistral AI vs. Silicon Valley: The Rise of Sovereign AI
TBPN Assertion Partly supported
Kilpatrick: Multimodal Foundation Models Match or Beat Domain-Specific Vision Models
“Relative to today where you can literally just write a prompt and send images or videos to the model and have it do those tasks like with basically, you know, near or better accuracy than you would get from domain specific models is absolutely fascinating.”
Logan Kilpatrick Apr 25, 2025 ▶ 16:03 Google's AI Comeback in Their Own Words - Logan Kilpatrick
Nguyen: Pixel-based perception is much harder to scale than language in AI
“Much of it is, like because right now the models operating on, like, pixels instead of, like, language or whatnot, like, pixels is actually really, really hard for the models because, like, perception or visual perception. I think there's still, like, a lot of…”
Karina Nguyen Feb 9, 2025 ▶ 1:10:12 OpenAI researcher on why soft skills are the future of work | Karina Nguyen
LATENT SPACE Prediction Not checkable as stated
Tay: Multimodal AI architectures will eventually move completely to early fusion
“As early fusion models get more traction, I think the themes will start to get more and more, like, it's a bit like how all the tasks like unify, like from Like, two zero one nine to, like, now it's like all the tasks are unifying, now it's like all the modali…”
Yi Tay Jul 5, 2024 ▶ 1:31:55 The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
NO PRIORS Insight
Gu: High-quality speech synthesis requires multimodal foundation models
“And so actually to really get, like, perfect even just TTS or, like, speech-to-speech you actually really need to have, like, a model that has, More understanding, like at least of the language, but kind of like, it's not really an isolated component anymore. …”
Albert Gu Jun 27, 2024 ▶ 24:27 No Priors Ep. 70 | With Cartesia Co-Founders Karan Goel & Albert Gu
NO PRIORS Disclosure
Goel: Cartesia intends to train an on-device multimodal model
“Yes. But, you know, we have our own sort of set of techniques that we're developing in order to be able to do that effectively. I think that I will maybe leave for another podcast. But I think, yeah, I think that is the intention at the end of the day is build…”
Karan Goel Jun 27, 2024 ▶ 27:54 No Priors Ep. 70 | With Cartesia Co-Founders Karan Goel & Albert Gu
Albrecht: Vision is not essential for most coding and reasoning agent tasks
“And actually we found that for most of the kind of like code writing and reasoning problems that we care about, the visual part isn't really a huge important part of it.”
Josh Albrecht Jun 25, 2024 ▶ 45:17 State of the Art: Training 70B LLMs on 10,000 H100 clusters
LATENT SPACE Prediction Not checkable as stated
Luan: AI will converge into a universal byte model across all modalities
“Multimodal models are becoming more of a thing, we're behavioral cloning the visual world, but really what we're just going to have is this like universal byte model, right? Where like tokens of data that have high signal come in, and then all of those pattern…”
David Luan Mar 27, 2024 ▶ 6:15 Why Google failed to make GPT-3 -- with David Luan of Adept
LATENT SPACE Prediction Held up
Multimodal models will completely supplant text-only large language models
“I actually think like it's really clear today. Multimodal models are the default foundation model, right? It's just going to supplant LLMs. Like why did you just train a giant multimodal model?”
David Luan Mar 27, 2024 ▶ 28:05 Why Google failed to make GPT-3 -- with David Luan of Adept
Ruiz: Vision models effectively interpret mixed visual assets on infinite canvases
“The fun of the Infinite Canvas and Teal Draw in particular is that you could just dump like whatever you want onto the canvas. Screenshots, text, images, other websites sticky notes, all that stuff. And the model, even as something that was in preview, like th…”
Steve Ruiz Jan 5, 2024 ▶ 42:52 The Accidental AI Canvas - with Steve Ruiz of tldraw
BIG TECHNOLOGY Prediction Not checkable as stated
Ramaswamy: Multimodal AI models will see strong adoption for parsing PDFs
“I think multimodal models will have a lot more business use cases where you're looking at PDFs.”
Sridhar Ramaswamy Sep 6, 2023 ▶ 30:42 Google’s Weird Year + Neeva Goes to Snowflake — With Sridhar Ramaswamy

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.