Multimodal LLM
topic on 3 shows · 4 statements across 4 episodes
Latent Space
Invest Like the Best
the a16z Podcast
4 statements about Multimodal LLM, every show
Levine: Adapting Multimodal LLMs to Robot Control Is a Key Advance
“I do think that the advent of Multimodal LLMs that can be adapted to robotic control to bring in that common sense. I do think that's a really important advance.”
Ben-Smith: End-to-end multimodal LLMs will replace modular audio transcription pipelines
“In the future that would just be put everything into a big multimodal LLM. And it will output everything that you want.”
Biilmann: Headless browsers won't remain web agent standards long-term
“For the web, there are questions of, like, how can we simplify The way agents get, like interact with the web, where right now the state of the art is kind of like a headless browser and taking screenshots and having a multimodal LLM take action and that, righ…”
Justin Johnson: Multimodal LLMs shoehorn visual data into 1D token sequences
“And now the multimodal LLMs that we're seeing now, you kind of end up shoehorning the other modalities into this underlying representation of a one D sequence of tokens. Now when we move to spatial intelligence, it's kind of going the other way. Where we're sa…”