Multimodal Language Models
topic on 2 shows · 2 statements across 2 episodes
2 statements about Multimodal Language Models, every show
Levine: Multimodal LLMs hold broad knowledge but lack physical grounding
“Multimodal language models are really good at pulling in knowledge and trying to articulate that knowledge. They're not very good at, like, grounding that knowledge in physical situations, but they know stuff.”
Sergey Brin: Pre-multimodal robotics efforts feel 'silly' in hindsight
“It, yeah, it just feels kind of silly having done all of that work and seeing now how capable these general language models are that include, for example, vision and image, and they're multimodal, and they can understand The scene and everything, and not havin…”